July 2026Unreviewed
What AI Red-Team Evaluations Can and Cannot Prove
Bandana Kaur
Abstract
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated e
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{kaur2026what,
title = {{What AI Red-Team Evaluations Can and Cannot Prove}},
author = {Bandana Kaur},
year = {2026},
month = jul,
eprint = {2607.21735},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/5ac1a19d9e04829701552cea010f3dbdd72e65f3}
}