Skip to content
Search
paperMay 2026Unreviewed

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard

Sahar Abdelnabi, Chris Hicks, Konrad Rieck, Ahmad-Reza Sadeghi

Abstract

The benchmarks used to evaluate AI agents in security-critical roles suffer from crucial weaknesses. Building on recent empirical evidence, we characterize three core challenges that undermine security evaluations: benchmark vulnerabilities, temporal staleness, and runtime uncertainty. We then outline practical directions toward building more robust and trustworthy evaluation frameworks.

Categories

Cite

@misc{abdelnabi2026measuring,
  title = {{Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard}},
  author = {Sahar Abdelnabi and Chris Hicks and Konrad Rieck and Ahmad-Reza Sadeghi},
  year = {2026},
  month = may,
  eprint = {2605.22568},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.22568}
}