Skip to content
Search
paperSeptember 2026Unreviewed

What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark

Seyed Mosayeb Alam

Abstract

Ask a language model the same question about the same document twenty times, and it sometimes returns two different answers. We built Probity, a benchmark of 60 tasks and 470 items from real venture-financing filings, to measure how often this happens. Then we audited our own corpus and found a defect any excerpt-built benchmark can carry: items whose evidence is missing from the window of text the model is shown. The audit flags 36 items and separates two failures a single flag would conflate:

Categories

Cite

@misc{alam2026what,
  title = {{What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark}},
  author = {Seyed Mosayeb Alam},
  year = {2026},
  month = sep,
  eprint = {2609.06147},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06147}
}