Allen Institute for AI (AI2) has introduced BenchMIRT, a new methodology that audits large language model benchmarks at the granular level of individual questions and tasks. Using multidimensional Item Response Theory—a technique borrowed from psychometrics—BenchMIRT analyzes how 100 different LLMs performed across 16 benchmarks comprising more than 34,000 questions to identify what capabilities are actually being measured.
Applied to six general reasoning benchmarks and ten safety-focused evaluations, BenchMIRT independently discovered that models' benchmark performance breaks down into two dominant capability dimensions: safety and general reasoning. However, the analysis revealed that many individual benchmarks measure more than their stated purpose. For example, BBQ, designed to test social bias, actually correlates more strongly with general reasoning ability than safety behavior. Similarly, WMDP, which evaluates dangerous dual-use knowledge, measures general reasoning more than safety—and interestingly, stronger reasoning associates with lower WMDP scores since refusal is the desired response.
The research shows that single benchmark scores often combine multiple confounded signals, potentially misleading researchers about true model capabilities. HarmBench demonstrates that while harmful content and contextual questions align with safety, copyright-related questions measure general reasoning instead. BenchMIRT's methodology enables researchers to disentangle these mixed signals and gain clearer insight into what drives benchmark performance, offering a more nuanced approach to LLM evaluation.
Key Points
BenchMIRT applies multidimensional Item Response Theory to identify what individual benchmark questions actually measure across LLM populations
Analysis of 100 LLMs across 16 benchmarks independently discovered two dominant capability dimensions: safety and general reasoning
Many benchmarks measure multiple capabilities simultaneously—BBQ correlates more with general reasoning than intended safety focus, WMDP correlates with reasoning not safety
Single benchmark scores often obscure mixed signals, creating misunderstandings about model capabilities and requiring better decomposition methods