June 2024ReviewedOpen access
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, Florian Tramer
arXiv preprint
Abstract
Introduces AgentDojo, a framework for evaluating the security of LLM agents against prompt injection and other attacks in realistic tool-use scenarios.
Categories
#agent-security#evaluation-framework#tool-use
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM06Excessive Agency
OWASP Top 10 for Agentic Applications
- ASI02Tool Misuse & Exploitation
- ASI01Agent Goal Hijack
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0053AI Agent Tool Invocation
Cite
@misc{debenedetti2024agentdojo,
title = {{AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents}},
author = {Edoardo Debenedetti and Jie Zhang and Mislav Balunovic and Luca Beurer-Kellner and Marc Fischer and Florian Tramer},
year = {2024},
month = jun,
eprint = {2406.13352},
archivePrefix = {arXiv},
doi = {10.52202/079017-2636},
url = {https://arxiv.org/abs/2406.13352}
}