Skip to content
Search
paperJuly 2026Unreviewed

What AI Red-Team Evaluations Can and Cannot Prove

Bandana Kaur

Abstract

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated e

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{kaur2026what,
  title = {{What AI Red-Team Evaluations Can and Cannot Prove}},
  author = {Bandana Kaur},
  year = {2026},
  month = jul,
  eprint = {2607.21735},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/5ac1a19d9e04829701552cea010f3dbdd72e65f3}
}