July 2026Unreviewed
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko
Abstract
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel opt
Categories
Cite
@misc{libon2026researcharena,
title = {{ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R\&D}},
author = {Lena Libon and Ben Rank and Jehyeok Yeon and David Schmotz and Jeremy Qin and Daniel Donnelly and Derck Prinzhorn and Maksym Andriushchenko},
year = {2026},
month = jul,
eprint = {2607.19321},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.19321}
}