Skip to content
Search
paperMay 2026Unreviewed

Stateful Online Monitoring Catches Distributed Agent Attacks

Davis Brown, Samarth Bhargav, Arav Santhanam, Kasper Hong, Ivan Zhang, Matan Shtepel, Steffi Chern, Alexander Robey, Eric Wong, Hamed Hassani

Abstract

Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection, attackers frequently distribute their misuse, splitting a harmful task across many user accounts so each individual transcript looks benign. Because safety monitors score only one agent context at a time, they are structurally blind to misuse that is only visible in aggregate, across many accounts. We show this gap is real by building, to our know

Categories

Cite

@misc{brown2026stateful,
  title = {{Stateful Online Monitoring Catches Distributed Agent Attacks}},
  author = {Davis Brown and Samarth Bhargav and Arav Santhanam and Kasper Hong and Ivan Zhang and Matan Shtepel and Steffi Chern and Alexander Robey and Eric Wong and Hamed Hassani},
  year = {2026},
  month = may,
  eprint = {2605.31593},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.31593}
}