June 2025Unreviewed
STACK: Adversarial Attacks on LLM Safeguard Pipelines
Ian R. McKenzie, O. Hollinsworth, Tom Tseng, Xander Davies, Stephen Casper, A. D. Tucker, Robert Kirk, Adam Gleave
AAAI Conference on Artificial Intelligence
Abstract
Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic guards their latest Claude 4 Opus model using one such defense pipeline, and other frontier developers including Google DeepMind and OpenAI pledge to soon deploy similar defenses. However, the security of such pipelines is unclear, with limited prior work evaluating or attacking these pipelines. We address this gap by developing and red-teaming an open-source defense pipelin
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{mckenzie2025stack,
title = {{STACK: Adversarial Attacks on LLM Safeguard Pipelines}},
author = {Ian R. McKenzie and O. Hollinsworth and Tom Tseng and Xander Davies and Stephen Casper and A. D. Tucker and Robert Kirk and Adam Gleave},
year = {2025},
month = jun,
booktitle = {AAAI Conference on Artificial Intelligence},
eprint = {2506.24068},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2506.24068},
url = {https://www.semanticscholar.org/paper/59db5498605113c8b2567219a5041ff08f7a4587}
}