June 2026Unreviewed
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad, Joachim Schaeffer, Ram Potham, Tyler Tracy
Abstract
An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but untrusted AI agents under the oversight of a weaker, trusted monitor and a limited human audit budget. Control evaluations stress-test these protocols by pitting a red-team attack policy against the blue-team monitor, but current evaluations typically assume attackers that do not strategically select when to attack. We st
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{gewang2026attack,
title = {{Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety}},
author = {Catherine Ge-Wang and Tyler Crosse and Benjamin Hadad and Joachim Schaeffer and Ram Potham and Tyler Tracy},
year = {2026},
month = jun,
eprint = {2606.06529},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.06529}
}