ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
By Lena Libon · Paper · cs.AI
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a moni