activity
20242026
collaborators

9 papers

cs.LG2026

Efficiently Aligning Language Models with Online Natural Language Feedback

Christine Ye, Joe Benton

Reinforcement learning with verifiable rewards has been used to elicit impressive performance from language models in many domains. But, broadly beneficial deployments of AI may re…

cs.AI2025

Inverse Scaling in Test-Time Compute

Aryo Pradipta Gema, Alexander Hägele, Runjin Chen +11

We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between tes…

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI2025

Natural Emergent Misalignment from Reward Hacking in Production RL

Monte MacDiarmid, Benjamin Wright, Jonathan Uesato +19

We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment. We start with a pretrained model, im…

cs.CL2025

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Miles Turpin, Andy Arditi, Marvin Li +2

Language models trained with reinforcement learning (RL) can engage in reward hacking--the exploitation of unintended strategies for high reward--without revealing this behavior in…

cs.AI2025

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

Jonathan Kutasov, Yuqi Sun, Paul Colognese +9

As Large Language Models (LLMs) are increasingly deployed as autonomous agents in complex and long horizon settings, it is critical to evaluate their ability to sabotage users by p…