Skip to content
Search
paperJune 2026Unreviewed

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

Kargi Chauhan, Pratibha Revankar

Abstract

LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt injection to induce credential exfiltration. We study this failure mode through three complementary defenses. First, we ask whether activation probes can detect credential access before output tokens are emitted. Second, we construct honeytokens from format-specific character models and calibrate detection with split conformal prediction. Third, we tr

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{chauhan2026caught,
  title = {{Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents}},
  author = {Kargi Chauhan and Pratibha Revankar},
  year = {2026},
  month = jun,
  eprint = {2606.04141},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.04141}
}