2 papers
cs.AI2026
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon +5
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such…
cs.LG2025
Expressive Reward Synthesis with the Runtime Monitoring Language
Daniel Donnelly, Angelo Ferrando, Francesco Belardinelli
A key challenge in reinforcement learning (RL) is reward (mis)specification, whereby imprecisely defined reward functions can result in unintended, possibly harmful, behaviours. In…