September 2026Unreviewed
LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs
Muhammed Saeed, Simon Razniewski
Abstract
Flagship language models appear saturated on benchmarks like MMLU (Hendrycks et al., 2021), scoring above 90% - yet benchmarks test only what the experimenter thought to ask, the availability bias of fixed question sets. LLMPEDIA makes this bias measurable and browsable. We recursively materialized ~1.3M articles from three model families' parametric memory (GPT-5-mini, DeepSeek-V3.2, Llama-3.3-70B) without retrieval, then audited a stratified sample of atomic claims against Wikipedia and a cura
Categories
Cite
@misc{saeed2026llmpedia,
title = {{LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs}},
author = {Muhammed Saeed and Simon Razniewski},
year = {2026},
month = sep,
eprint = {2609.01182},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01182}
}