Skip to content
Search
paperJune 2024ReviewedOpen access

AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents

Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, Florian Tramer

arXiv preprint

Abstract

Introduces AgentDojo, a framework for evaluating the security of LLM agents against prompt injection and other attacks in realistic tool-use scenarios.

Categories

#agent-security#evaluation-framework#tool-use

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM06Excessive Agency
OWASP Top 10 for Agentic Applications
  • ASI02Tool Misuse & Exploitation
  • ASI01Agent Goal Hijack
MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0053AI Agent Tool Invocation

Cite

@misc{debenedetti2024agentdojo,
  title = {{AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents}},
  author = {Edoardo Debenedetti and Jie Zhang and Mislav Balunovic and Luca Beurer-Kellner and Marc Fischer and Florian Tramer},
  year = {2024},
  month = jun,
  eprint = {2406.13352},
  archivePrefix = {arXiv},
  doi = {10.52202/079017-2636},
  url = {https://arxiv.org/abs/2406.13352}
}