September 2026Unreviewed
What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark
Seyed Mosayeb Alam
Abstract
Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built benchmark can carry: items whose evidence is missing from the window of text the model is shown. The audit flags 36 items and separates two failures a single flag would conflate:
Categories
Cite
@misc{alam2026what,
title = {{What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark}},
author = {Seyed Mosayeb Alam},
year = {2026},
month = sep,
eprint = {2609.06147},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.06147}
}