Skip to content
Search
paperSeptember 2026Unreviewed

What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores

Dana Paquin, Riddhiman Jain

Abstract

Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extr

Categories

Cite

@misc{paquin2026what,
  title = {{What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores}},
  author = {Dana Paquin and Riddhiman Jain},
  year = {2026},
  month = sep,
  eprint = {2609.09372},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.09372}
}