May 2026Unreviewed
A New Framework for Cybersecurity Refusals in AI Agents
Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson, J Zico Kolter
Abstract
Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domains like cybersecurity. Existing benchmarks for AI agents in cybersecurity focus mainly on measuring proficiency--how effectively agents can complete offensive security tasks--but neglect a critical question: when and how should agents refuse harmful requests? We present the first framework for establishing refusal boundaries in offensive security c
Categories
Cite
@misc{jones2026new,
title = {{A New Framework for Cybersecurity Refusals in AI Agents}},
author = {Eliot Krzysztof Jones and Mateusz Dziemian and Matt Fredrikson and J Zico Kolter},
year = {2026},
month = may,
eprint = {2606.02644},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.02644}
}