Skip to content
Search
paperJuly 2026Unreviewed

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko

Abstract

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel opt

Categories

Cite

@misc{libon2026researcharena,
  title = {{ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R\&D}},
  author = {Lena Libon and Ben Rank and Jehyeok Yeon and David Schmotz and Jeremy Qin and Daniel Donnelly and Derck Prinzhorn and Maksym Andriushchenko},
  year = {2026},
  month = jul,
  eprint = {2607.19321},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.19321}
}