Skip to content
Search
paperJune 2026Unreviewed

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le

Abstract

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{he2026pihunter,
  title = {{PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections}},
  author = {Pengfei He and Lesly Miculicich and Vishesh Sharma and Ash Fox and George Lee and Jiliang Tang and Tomas Pfister and Long T. Le},
  year = {2026},
  month = jun,
  eprint = {2606.12737},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.12737}
}