Skip to content
Search
paperAugust 2026Unreviewed

Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection

Wu-Jie Xiong, Rabimba Karanjai, Yang Lu, W. Shi, Lei Xu

Abstract

Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to d

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{xiong2026reachabilitybased,
  title = {{Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection}},
  author = {Wu-Jie Xiong and Rabimba Karanjai and Yang Lu and W. Shi and Lei Xu},
  year = {2026},
  month = aug,
  eprint = {2608.30041},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/ca9874d087d8df198b302c54402b0df91d19bbbc}
}