June 2026Unreviewed
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
Reza Soosahabi, Vivek Namsani
Abstract
Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation. This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{soosahabi2026analyzing,
title = {{Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems}},
author = {Reza Soosahabi and Vivek Namsani},
year = {2026},
month = jun,
eprint = {2606.20470},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.20470}
}