August 2026Unreviewed
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye
Abstract
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering an
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{mou2026toolhazard,
title = {{ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents}},
author = {Yutao Mou and Pengfei Yang and Zhe Yin and Zhangchi Xue and Xiaotian Luan and Dingyao Yu and Tong Zhang and Shikun Zhang and Wei Ye},
year = {2026},
month = aug,
eprint = {2608.11878},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.11878}
}