Skip to content
Search
paperSeptember 2026Unreviewed

LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs

Muhammed Saeed, Simon Razniewski

Abstract

Flagship language models appear saturated on benchmarks like MMLU (Hendrycks et al., 2021), scoring above 90% - yet benchmarks test only what the experimenter thought to ask, the availability bias of fixed question sets. LLMPEDIA makes this bias measurable and browsable. We recursively materialized ~1.3M articles from three model families' parametric memory (GPT-5-mini, DeepSeek-V3.2, Llama-3.3-70B) without retrieval, then audited a stratified sample of atomic claims against Wikipedia and a cura

Categories

Cite

@misc{saeed2026llmpedia,
  title = {{LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs}},
  author = {Muhammed Saeed and Simon Razniewski},
  year = {2026},
  month = sep,
  eprint = {2609.01182},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01182}
}