papers

Publications (10)

cs.LG2026

Modular Pretraining Enables Access Control

Ethan Roland, Murat Cubuktepe, Erick Martinez +8

AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limi…

cs.AI2025

Natural Emergent Misalignment from Reward Hacking in Production RL

Monte MacDiarmid, Benjamin Wright, Jonathan Uesato +19

We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment. We start with a pretrained model, im…

cs.LG2025

Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs

Igor Shilov, Alex Cloud, Aryo Pradipta Gema +5

Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenge…

cs.LG2025

Output Supervision Can Obfuscate the Chain of Thought

Jacob Drori, Luke Marks, Bryce Woodworth +2

OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep C…

cs.GT2022

Anticipatory Fictitious Play

Alex Cloud, Albert Wang, Wesley Kerr

Fictitious play is an algorithm for computing Nash equilibria of matrix games. Recently, machine learning variants of fictitious play have been successfully applied to complicated…

cs.LG2025

Distillation Robustifies Unlearning

Bruce W. Lee, Addie Foote, Alex Infanger +6

Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: t…

cs.LG2025

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

Alex Cloud, Minh Le, James Chua +5

We study subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model w…

stat.OT2020

Variance decompositions for extensive-form games

Alex Cloud, Eric Laber

Quantitative measures of randomness in games are useful for game design and have implications for gambling law. We treat the outcome of a game as a random variable and derive a clo…

cs.AI2026

Recontextualization Mitigates Specification Gaming without Modifying the Specification

Ariana Azarbal, Victor Gillioz, Vladimir Ivanov +6

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game"…

cs.LG2024

Gradient Routing: Masking Gradients to Localize Computation in Neural Networks

Alex Cloud, Jacob Goldman-Wetzler, Evžen Wybitul +2

Neural networks are trained primarily based on their inputs and outputs, without regard for their internal mechanisms. These neglected mechanisms determine properties that are crit…