Skip to content
Search
paperJuly 2026Unreviewed

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

Oliver Makins, Orazio Angelini, Zohreh Shams, Mary Phuong

Abstract

AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single agent in one trajectory, but real deployments run many agents over shared infrastructure, and the most severe risks (model-weight exfiltration, training-run poisoning) plausibly need several agents acting in concert. We initiate the empirical study of multi-agent AI control, formalising distributed attacks in which several agents jointly aim for a

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{makins2026multiagent,
  title = {{Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors}},
  author = {Oliver Makins and Orazio Angelini and Zohreh Shams and Mary Phuong},
  year = {2026},
  month = jul,
  eprint = {2607.07368},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.07368}
}