Skip to content
Search
paperAugust 2026Unreviewed

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

William Caban

Abstract

Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, V_total <= V_1 x V_2 x V_3, that captures multiplicative degradation across task generation (V_1), human-simulator calibration (V_2), and automated judgment (V_3). Under empirically grounded estim

Categories

Cite

@misc{caban2026measurement,
  title = {{Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation}},
  author = {William Caban},
  year = {2026},
  month = aug,
  eprint = {2608.00794},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.00794}
}