July 2026Unreviewed
Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
Oliver Makins, Orazio Angelini, Zohreh Shams, Mary Phuong
Abstract
AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single agent in one trajectory, but real deployments run many agents over shared infrastructure, and the most severe risks (model-weight exfiltration, training-run poisoning) plausibly need several agents acting in concert. We initiate the empirical study of multi-agent AI control, formalising distributed attacks in which several agents jointly aim for a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{makins2026multiagent,
title = {{Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors}},
author = {Oliver Makins and Orazio Angelini and Zohreh Shams and Mary Phuong},
year = {2026},
month = jul,
eprint = {2607.07368},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.07368}
}