August 2026Unreviewed
Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection
Wu-Jie Xiong, Rabimba Karanjai, Yang Lu, W. Shi, Lei Xu
Abstract
Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed operations. They do not directly address how an agent's future authority should change once untrusted data enters its state. We present SkillGuard, a harness-level enforcement layer that treats this event as contamination and restricts future capabilities to d
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{xiong2026reachabilitybased,
title = {{Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection}},
author = {Wu-Jie Xiong and Rabimba Karanjai and Yang Lu and W. Shi and Lei Xu},
year = {2026},
month = aug,
eprint = {2608.30041},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/ca9874d087d8df198b302c54402b0df91d19bbbc}
}