The UK AI Security Institute has released standardized evaluation results for frontier large language models in partnership with the EvalEval Coalition, marking progress toward more transparent and reproducible AI benchmarking. The release includes verified performance data from five major benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0—across six frontier models from Anthropic (Claude Opus 4, 4.5, 4.6) and OpenAI (GPT-5, 5.2, 5.4). The results accompany a peer-reviewed paper examining how inference-time compute affects model performance and are shared through EvalEval's standardized "Every Eval Ever" schema and "Evaluation Cards" platform.
The collaboration addresses a fundamental gap in AI evaluation practice: benchmark results are typically reported across disparate formats with insufficient detail for reproduction or comparison. By adopting a shared schema and making transcript-level data and evaluation configurations publicly available, AISI and EvalEval create reference points that help researchers understand how different setup choices—such as whether models receive feedback during testing—influence reported performance. This transparency is particularly important as evaluations become primary evidence sources for understanding model capabilities and limitations.
The release represents one of the first large-scale adoptions of the EvalEval Coalition's standardized approach, potentially establishing a template for how research organizations and AI developers will share evaluation data in the future. The initiative aligns with growing emphasis on evaluation rigor and reproducibility in the AI research community, with implications for how model developers, evaluation researchers, and policymakers assess frontier AI systems.
Key Points
UK AISI releases verified benchmark results using standardized EvalEval schema, covering 5 major benchmarks and 6 frontier models
Release provides transcript-level transparency and configuration details to enable reproduction and meaningful comparison of evaluations
Standardized approach addresses reproducibility gaps that currently plague AI benchmark reporting across the industry
Represents first large-scale adoption of Every Eval Ever schema, potentially establishing industry template for evaluation data sharing