What UK AISI And EvalEval Bring To AI Benchmark Reproducibility
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What UK AISI And EvalEval Bring To AI Benchmark Reproducibility on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, covering five benchmarks across six frontier models plus two cyber evaluations. The release ties results to setup details because the accompanying AISI paper shows scores shift with inference-time compute and evaluation protocol.

The UK AI Security Institute (AISI) has begun publishing selected AI benchmark results through EvalEval’s Evaluation Cards, pairing scores with verification, context and configuration details that show how each result was produced. The release covers five benchmarks across six frontier models and accompanies AISI’s paper on how inference-time compute changes evaluation outcomes.

The main experiment in the release covers results for HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The models included are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4.

AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones. Those use a different set of models that overlaps only partly with the main experiment, so the same model list should not be assumed to apply to them, and the announcement does not enumerate that set.

The records are tied to AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how scores depend on inference-time compute and evaluation protocol. For Humanity’s Last Exam, the analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.

EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI says publicly reported material is being made available “where appropriate” — the announcement does not claim every AISI evaluation or underlying transcript is included.

At a glance
announcementWhen: announced alongside AISI’s inference-co…
The developmentUK AISI has released evaluation records on EvalEval’s Evaluation Cards platform, alongside its paper on how inference-time compute shapes evaluation results.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Setup Details Change Score Comparisons

Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis illustrates the problem: outcomes shifted with inference compute and with whether models received correctness feedback between attempts. A score reported without those conditions leaves readers unsure what performance it actually represents.

Publishing results with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that treats evaluations as evidence about advanced AI capabilities. The records do not settle which benchmark or protocol is best, but they make the conditions behind a result easier to see.

Amazon

Top picks for "aisi evaleval benchmark"

As an affiliate, we earn on qualifying purchases.

From NeurIPS Workshop to Shared Schema

The release builds on earlier collaboration between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that shared infrastructure to AISI’s publicly reported methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model metadata. Together, these efforts address a practical reporting problem: results published across formats and outlets may omit details needed to interpret or reproduce a run, while repeating costly evaluations is often not feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Coverage and Reproduction Limits

The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. Because it says publicly reported methods and findings are being made available “where appropriate,” the release should not be read as a complete archive of all AISI evaluation work.

The cyber evaluations use a different, partly overlapping model set that has not been enumerated. The announcement also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge current coverage and compare records consistently.

Broader Adoption of Every Eval Ever

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema.

Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of the records contributors publish. No further release date or adoption milestone was specified.

Key Questions

Which models and benchmarks are covered in the release?

The main experiment covers HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0, tested on Claude Opus 4, 4.5 and 4.6, plus GPT-5, GPT-5.2 and GPT-5.4.

Do the cyber evaluations use the same models?

No. The Cyber CTFs and The Last Ones evaluations use a different model set that overlaps only partly with the main experiment, and AISI has not enumerated it.

Does the release include all of AISI’s evaluation work?

No. AISI says publicly reported methods and findings are being made available “where appropriate.” The announcement does not claim complete coverage of every AISI evaluation or transcript.

Why does evaluation setup matter for benchmark scores?

AISI’s paper found that scores on Humanity’s Last Exam changed with inference-time compute and with whether models received correctness feedback between attempts, meaning the same model can score differently under different protocols.

What is Every Eval Ever?

It is EvalEval’s shared schema for documenting evaluations, developed with feedback from AISI following a joint workshop at NeurIPS 2025. It organizes benchmark metadata, run data and model metadata into a common format.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Strategic Significance Of Nvidia Acquiring The Open Commons For AI

Nvidia reportedly agrees to acquire Hugging Face for $12.9 billion, aiming to control open-source AI models and reinforce its market dominance amid regulatory concerns.

Rethinking AI Bottlenecks: Moving Beyond Model Improvements

New insights reveal that AI deployment bottlenecks now lie in system integration and orchestration, favoring small operators owning their entire stack.

When One Agent Isn’t Enough: Claude Now Builds Its Own Team Of Agents On The Fly

Claude now builds its own team of agents on the fly, enabling complex, high-value tasks through dynamic workflows that orchestrate multiple subagents.

Unlocking The Power Of AI With SeedRealtime: ByteDance Seed’s Innovative Multi-modal Full-duplex Model

ByteDance Seed announced SeedRealtime, a native audio-visual, full-duplex large language model capable of watching, listening, and speaking in real time.