Skip to content
Search
paperJuly 2026Unreviewed

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu

Abstract

Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run. Per-step safety checks that judge each action in isolation may fail to recognize the complete distributed payload. We investigate how early such an attack can be detected while the run is still unfolding, and how robustly it can be caught on

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{arias2026early,
  title = {{Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study}},
  author = {Diego Fernandez Arias and Dev Prashant Mistry and Ren Wang and Yibo Hu},
  year = {2026},
  month = jul,
  eprint = {2607.24893},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.24893}
}