Skip to content
Search
paperJune 2026Unreviewed

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems

Reza Soosahabi, Vivek Namsani

Abstract

Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{soosahabi2026analyzing,
  title = {{Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems}},
  author = {Reza Soosahabi and Vivek Namsani},
  year = {2026},
  month = jun,
  eprint = {2606.20470},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.20470}
}