Skip to content
Search
paperAugust 2026Unreviewed

The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents

Md Habibur Rahman, Jaeho Kim

Abstract

A tool-using LLM agent that reads attacker-controlled web content while holding a secret faces indirect prompt injection: the content may make it exfiltrate the secret. In a safe synthetic lab (canary secret, mock tools, matched clean-vs-poisoned metric) we report the framing gap: across six models, ten overt injection classes are refused (gpt-4o 0%), but reframing the identical leak as a mandatory integrity signature, config field, or look-alike "trusted" host drives gpt-4o 0% to 100%. The atta

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{rahman2026framing,
  title = {{The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents}},
  author = {Md Habibur Rahman and Jaeho Kim},
  year = {2026},
  month = aug,
  eprint = {2608.27092},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.27092}
}