🔍 Read the full analysis: How UK AISI And EvalEval Help Make AI Benchmark Results Reproducible on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected results from five benchmarks across six frontier models through EvalEval’s Evaluation Cards. The records add verification, evaluation context and configuration information, while the coverage of evaluations and model sets remains limited.
The UK AI Security Institute (AISI) is publishing selected results from its AI evaluations through EvalEval’s Evaluation Cards, which pair scores with verification, context and configuration information. The release covers five benchmarks across six frontier models, as well as two cyber evaluations, and accompanies an AISI paper examining how inference-time compute and evaluation protocols affect reported performance.
The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The reported model set for those experiments comprises Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results from Cyber CTFs and The Last Ones, two cyber evaluations that use a different, partly overlapping model set. The six models listed for the main experiment should not be taken as the model list for those cyber results.
The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The paper examines how evaluation results vary with the amount of compute used during inference and with the protocol applied. For Humanity’s Last Exam, its analysis tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.
EvalEval describes the cards as bringing together verified results, evaluation context and configuration information. Its platform uses a common format for benchmark metadata, evaluation-run data and model metadata. AISI’s public reporting is included where appropriate, but the announcement does not say that every evaluation or underlying transcript is part of the release.
Why Evaluation Conditions Matter
Benchmark scores are often compared across models, but the number alone may not show what conditions produced it. AISI’s analysis of Humanity’s Last Exam illustrates the issue: measured outcomes changed with inference compute and with whether a model received correctness feedback between attempts. Without those details, readers may have difficulty telling what a score represents or whether two reported results are comparable.
Publishing results alongside information about their setup gives researchers and practitioners a way to inspect individual runs and judge whether differences in scores could reflect different protocols. This can inform work in model research, development and policy, where benchmark results may be used as evidence about advanced AI capabilities. The cards do not determine which benchmark or protocol is best, and they do not establish that every published result can be independently reproduced. They make some of the conditions behind selected results easier to examine.
That distinction matters because repeating evaluations can be costly, while published results may appear in formats that leave out details needed to interpret them. A shared record can help readers find those details and compare runs under more consistent terms. Its usefulness still depends on what contributors provide: common formatting cannot fill in missing information or make results comparable when their methods differ.
The release builds on work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever, or EEE, a shared schema for documenting evaluations. The current records apply that infrastructure to methods and findings AISI has made public where appropriate.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas including transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring together results with information about benchmarks and models. These efforts address a reporting challenge: when findings are spread across papers and other formats, the conditions behind a result may be difficult to locate or compare.
The paper provides a concrete example of why those conditions matter. Its analysis does not treat a benchmark score as independent of the evaluation process; it examines how inference-time compute and feedback relate to outcomes. The released cards offer a way to connect selected reported results with relevant run information, while the paper sets out the analysis behind them.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
What the Published Records Cover
The announcement does not specify how many records or transcripts are available, which setup fields appear for every benchmark, or whether outside researchers have independently reproduced the results. Because AISI says material is being shared where appropriate, the release should not be read as a complete archive of its evaluation work.
The announcement also does not enumerate the model set used for Cyber CTFs and The Last Ones, beyond saying it differs from and partly overlaps with the main experiment’s set. It gives no record-by-record release dates and does not describe how disagreements between results produced under different protocols would be handled. Those details would help readers assess coverage and make consistent comparisons across the collection.
It also remains unclear how consistently contributors will supply the information needed to interpret each result. A shared schema provides a common structure, but the material available on each card may depend on the reporting and records provided for that evaluation.
Broader Use of the EEE Schema
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Its stated next step is broader adoption of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmark and run data using the schema. Researchers working in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection.
Wider participation could make it easier to find and compare evaluation records, provided contributors publish sufficiently complete and consistent details. Neither a further release date nor an adoption milestone was specified, so the timing and scale of any expansion remain unknown.
Key Questions
What has AISI released?
AISI has made selected AI benchmark results available through EvalEval’s Evaluation Cards, with verification, context and configuration information. The release also includes results from two cyber evaluations.
Which models are covered by the five main benchmarks?
The reported set is Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The announcement says the cyber evaluations use a different, partly overlapping set, which it does not enumerate.
Why do inference compute and feedback matter?
AISI’s paper examines how evaluation outcomes can depend on these conditions. In its Humanity’s Last Exam analysis, models given correctness feedback after attempts went on to solve additional tasks as token use increased.
Does the release include every AISI evaluation?
No such claim was made. AISI’s public reporting is included where appropriate, and the announcement does not describe the release as a complete archive of all evaluations or transcripts.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
