5 papers
Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization
Frank Xiao, Mary Phuong
Model post-training, and in particular reinforcement learning (RL), is one of the primary mechanisms by which developers can shape models' values and behaviors. However, as models…
Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents
Frank Xiao, Mary Phuong
Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render tr…
Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
Frank Xiao, Santiago Aranguri
We propose probe-based data attribution, a method that traces behavioral changes in post-trained language models to responsible training datapoints. By computing activation-differe…
Why Do Language Model Agents Whistleblow?
Kushal Agrawal, Frank Xiao, Guido Bergman +1
The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in…
SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
Gio Huh, Dhruv Sheth, Rayhan Zirvi +1
While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smal…