Skip to content
Search
paperSeptember 2026Unreviewed

AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories

Asif Pinjari, Mithun Paul Saint-Germain

Abstract

LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read in order: a benign prefix gives way to actions that serve the attacker rather than the user. Existing benchmarks measure whether such attacks succeed against live agents, and existing guard models judge a trace as a whole; no public corpus labels, step by step,

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{pinjari2026agentdrift,
  title = {{AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories}},
  author = {Asif Pinjari and Mithun Paul Saint-Germain},
  year = {2026},
  month = sep,
  eprint = {2609.06972},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06972}
}