August 2026Unreviewed
CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?
Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu
Abstract
Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known injection attacks, deploying them in LLM agent environments remains challenging due to attack variants and emerging threats. Moreover, existing solutions typically suffer from an inherent trilemma, i.e., a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{liang2026caitlyn,
title = {{CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?}},
author = {Zi Liang and Xiaoyu Xu and Yanyun Wang and Minxin Du and Qingqing Ye and Haibo Hu},
year = {2026},
month = aug,
eprint = {2608.27990},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.27990}
}