May 2026Unreviewed
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injectio
Lei Zhao, Abhay Bhaskar, Edgar Dobriban
Abstract
AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) risk: an agent may execute harmful instructions embedded in untrusted inputs such as email, downloaded files, webpages, repositories, or group-chat messages. Existing evaluations are often small, purely simulated, or focused on a narrow set of channels. We introduce LivePI (Live Prompt Injection), a structured benchmark for IPI risk in a production-l
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{zhao2026livepi,
title = {{LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injectio}},
author = {Lei Zhao and Abhay Bhaskar and Edgar Dobriban},
year = {2026},
month = may,
eprint = {2605.17986},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.17986}
}