June 2026Unreviewed
PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections
Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le
Abstract
Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources. Existing defenses mainly focus on blocking malicious content at inference time, and current red-teaming methods primarily optimize attack success. As a result, developers have limited visibility into how latent prompt injections emerge and propagate through agents. We
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{he2026pihunter,
title = {{PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections}},
author = {Pengfei He and Lesly Miculicich and Vishesh Sharma and Ash Fox and George Lee and Jiliang Tang and Tomas Pfister and Long T. Le},
year = {2026},
month = jun,
eprint = {2606.12737},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.12737}
}