5 papers
Activation Steering via Generative Causal Mediation
Aruna Sankaranarayanan, Amir Zur, Atticus Geiger +1
Where should we intervene in a language model (LM) to localize and control behaviors that are diffused across many tokens of a long-form response? We introduce Generative Causal Me…
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
Prajna Soni, Deepika Raman, Dylan Hadfield-Menell
Datasets play a central role in AI governance by enabling both evaluation (measuring capabilities) and alignment (enforcing values) along axes such as helpfulness, harmlessness, to…
Layered Unlearning for Adversarial Relearning
Timothy Qian, Vinith Suriyakumar, Ashia Wilson +1
Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations. We are particularly interes…
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
Aruna Sankaranarayanan, Dylan Hadfield-Menell, Aaron Mueller
All natural languages are structured hierarchically. In humans, this structural restriction is neurologically coded: when two grammars are presented with identical vocabularies, br…
Goal Inference from Open-Ended Dialog
Rachel Ma, Jingyi Qu, Andreea Bobu +1
Embodied AI Agents are quickly becoming important and common tools in society. These embodied agents should be able to learn about and accomplish a wide range of user goals and pre…