May 2026Unreviewed
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Yongxiang Li, Moxin Li, Zhixin Ma, Fengbin Zhu, Dongrui Liu, Wenjie Wang, Fuli Feng
Abstract
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful behavior within one user request. However, we show tha
Categories
Cite
@misc{li2026plant,
title = {{Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents}},
author = {Yongxiang Li and Moxin Li and Zhixin Ma and Fengbin Zhu and Dongrui Liu and Wenjie Wang and Fuli Feng},
year = {2026},
month = may,
eprint = {2605.28201},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.28201}
}