July 2026Unreviewed
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
Abstract
Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run. Per-step safety checks that judge each action in isolation may fail to recognize the complete distributed payload. We investigate how early such an attack can be detected while the run is still unfolding, and how robustly it can be caught on
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{arias2026early,
title = {{Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study}},
author = {Diego Fernandez Arias and Dev Prashant Mistry and Ren Wang and Yibo Hu},
year = {2026},
month = jul,
eprint = {2607.24893},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.24893}
}