UK AISI shares benchmark results with setup details
Benchmark scores now arrive with the setup details needed to judge whether two results were measured the same way.
Published · on huggingface.co · 2 min read

The UK AI Security Institute is sharing evaluation methods and findings through EvalEval's Evaluation Cards, the EvalEval Coalition described on huggingface.co on 2026-09-22. The release covers a defined set of benchmarks and models, and it matters most to people who compare model scores across reports.
What AISI is sharing
The release covers five benchmarks from the main experiment of AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. Results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Two related cyber evaluations, Cyber CTFs and The Last Ones, are also included, using a different and partially overlapping set of models.
Each Evaluation Card carries verified results plus context and configuration information. That is the substance of the change. A score on its own says little about how it was produced, and the card is designed to record the conditions behind the reported figure.
The collaboration has a history. AISI and EvalEval began working together at a joint workshop alongside NeurIPS 2025, and feedback from the Institute shaped the Every Eval Ever schema. This phase puts that shared infrastructure into practice. AISI's related work in the same area includes OptStop, HiBayES, and standardisation efforts in transcript analysis and capability elicitation.
Why the configuration details matter
The paper studies how benchmark performance depends on inference-time compute and evaluation protocol. EvalEval's post illustrates this with Humanity's Last Exam: when models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased. That describes the reported experiment, not a general rule about model quality. It shows why the protocol belongs next to the number.
For teams that choose models, the practical value is comparison. Where a vendor reports a score without setup details, an Evaluation Card for the same benchmark and model offers a reference point for judging whether the two numbers are comparable at all. The cards can be browsed by benchmark or by model, so a reader can look up a specific claim rather than relying on a summary.
Limitations are worth stating plainly. Coverage is limited to the benchmarks and models listed above, and the cyber evaluations use a different model set. Evaluation Cards record what was run and how; they do not settle whether a benchmark measures what a buyer cares about. Reproducibility is a precondition for good comparison, not a substitute for it.
The useful habit is small: before repeating a benchmark number, check whether the configuration behind it is published. Where it is not, treat the figure as a claim rather than a measurement.
Source: huggingface.co — BARGO’s commentary on the linked source.
Source published: · Event date: