4 papers
Benchmarks for Detecting Measurement Tampering
Fabien Roger, Ryan Greenblatt, Max Nadeau +2
When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
Maximilian Li, Xander Davies, Max Nadeau
Language models often exhibit behaviors that improve performance on a pre-training objective but harm performance on downstream tasks. We propose a novel approach to removing undes…
Discovering Variable Binding Circuitry with Desiderata
Xander Davies, Max Nadeau, Nikhil Prakash +2
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-outp…
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi +29
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…