Skip to content
Search
paperAugust 2026Unreviewed

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

Sihan Hou, Xinmeng Hou, Zhijun Zhang, Zehao Wang, Xuhong Ren, Sibo Qin, Kuntharrgyal Khysru, Qing Guo

Abstract

Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external observations manipulate subsequent agent decisions and actions. Most existing adaptive attacks rely on repeatedly querying and refining against the target agent, whereas realistic attackers may have only a single opportunity to interact with an unknown target agent. We propose SAVOR (Strategy Abstraction Via Outcome-Conditioned Reflection), which sh

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{hou2026metacognitive,
  title = {{Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection}},
  author = {Sihan Hou and Xinmeng Hou and Zhijun Zhang and Zehao Wang and Xuhong Ren and Sibo Qin and Kuntharrgyal Khysru and Qing Guo},
  year = {2026},
  month = aug,
  eprint = {2608.08795},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.08795}
}