September 2026Unreviewed
What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores
Dana Paquin, Riddhiman Jain
Abstract
Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extr
Categories
Cite
@misc{paquin2026what,
title = {{What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores}},
author = {Dana Paquin and Riddhiman Jain},
year = {2026},
month = sep,
eprint = {2609.09372},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.09372}
}