Skip to content
Search
paperMay 2026Unreviewed

A New Framework for Cybersecurity Refusals in AI Agents

Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson, J Zico Kolter

Abstract

Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domains like cybersecurity. Existing benchmarks for AI agents in cybersecurity focus mainly on measuring proficiency--how effectively agents can complete offensive security tasks--but neglect a critical question: when and how should agents refuse harmful requests? We present the first framework for establishing refusal boundaries in offensive security c

Categories

Cite

@misc{jones2026new,
  title = {{A New Framework for Cybersecurity Refusals in AI Agents}},
  author = {Eliot Krzysztof Jones and Mateusz Dziemian and Matt Fredrikson and J Zico Kolter},
  year = {2026},
  month = may,
  eprint = {2606.02644},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.02644}
}