Skip to content
Search
paperJune 2026Unreviewed

AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents

Xinhang Ma, Taoran Li, Chaowei Xiao, Zhiyuan Yu, Ning Zhang, Yevgeniy Vorobeychik

Abstract

Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work have proposed a variety of defensive approaches against IPI. These can be grouped into three broad categories: 1) prompt-based (using prompting as a way to prevent agents from following malicious instructions), 2) detection-based (identifying and filtering malicious instructions), and 3) system-level (using systems insights, such as control and data isolation, for defense). However, com

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{ma2026autodojo,
  title = {{AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents}},
  author = {Xinhang Ma and Taoran Li and Chaowei Xiao and Zhiyuan Yu and Ning Zhang and Yevgeniy Vorobeychik},
  year = {2026},
  month = jun,
  eprint = {2606.15057},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.15057}
}