Skip to content
Search
paperJune 2025Unreviewed

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Ian R. McKenzie, O. Hollinsworth, Tom Tseng, Xander Davies, Stephen Casper, A. D. Tucker, Robert Kirk, Adam Gleave

AAAI Conference on Artificial Intelligence

Abstract

Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic guards their latest Claude 4 Opus model using one such defense pipeline, and other frontier developers including Google DeepMind and OpenAI pledge to soon deploy similar defenses. However, the security of such pipelines is unclear, with limited prior work evaluating or attacking these pipelines. We address this gap by developing and red-teaming an open-source defense pipelin

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@inproceedings{mckenzie2025stack,
  title = {{STACK: Adversarial Attacks on LLM Safeguard Pipelines}},
  author = {Ian R. McKenzie and O. Hollinsworth and Tom Tseng and Xander Davies and Stephen Casper and A. D. Tucker and Robert Kirk and Adam Gleave},
  year = {2025},
  month = jun,
  booktitle = {AAAI Conference on Artificial Intelligence},
  eprint = {2506.24068},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2506.24068},
  url = {https://www.semanticscholar.org/paper/59db5498605113c8b2567219a5041ff08f7a4587}
}