Skip to content
Search
paperAugust 2026Unreviewed

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu

Abstract

Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known injection attacks, deploying them in LLM agent environments remains challenging due to attack variants and emerging threats. Moreover, existing solutions typically suffer from an inherent trilemma, i.e., a

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{liang2026caitlyn,
  title = {{CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?}},
  author = {Zi Liang and Xiaoyu Xu and Yanyun Wang and Minxin Du and Qingqing Ye and Haibo Hu},
  year = {2026},
  month = aug,
  eprint = {2608.27990},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.27990}
}