activity
20242026
collaborators

5 papers

cs.CL2026

Activation Steering via Generative Causal Mediation

Aruna Sankaranarayanan, Amir Zur, Atticus Geiger +1

Where should we intervene in a language model (LM) to localize and control behaviors that are diffused across many tokens of a long-form response? We introduce Generative Causal Me…

cs.CY2025

CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment

Prajna Soni, Deepika Raman, Dylan Hadfield-Menell

Datasets play a central role in AI governance by enabling both evaluation (measuring capabilities) and alignment (enforcing values) along axes such as helpfulness, harmlessness, to…

cs.LG2025

Layered Unlearning for Adversarial Relearning

Timothy Qian, Vinith Suriyakumar, Ashia Wilson +1

Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations. We are particularly interes…

cs.CL2025

Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models

Aruna Sankaranarayanan, Dylan Hadfield-Menell, Aaron Mueller

All natural languages are structured hierarchically. In humans, this structural restriction is neurologically coded: when two grammars are presented with identical vocabularies, br…

cs.AI2024

Goal Inference from Open-Ended Dialog

Rachel Ma, Jingyi Qu, Andreea Bobu +1

Embodied AI Agents are quickly becoming important and common tools in society. These embodied agents should be able to learn about and accomplish a wide range of user goals and pre…