8 papers
A Dataset for Dynamic Human Preferences for Vision Language Models
Hannah Gao, Dylan Hadfield-Menell, Rachel Ma
Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models can adapt to real-time preferenc…
A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers
Hannah Gao, Isha Agarwal, Dylan Hadfield-Menell +1
The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances or perturbations, such as blurr…
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
Rachel Ma, Jingyi Qu, Andreea Bobu +1
We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferen…
Activation Steering via Generative Causal Mediation
Aruna Sankaranarayanan, Amir Zur, Atticus Geiger +1
Where should we intervene in a language model (LM) to localize and control behaviors that are diffused across many tokens of a long-form response? We introduce Generative Causal Me…
Goal Inference from Open-Ended Dialog
Rachel Ma, Jingyi Qu, Andreea Bobu +1
Embodied AI Agents are quickly becoming important and common tools in society. These embodied agents should be able to learn about and accomplish a wide range of user goals and pre…
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
Prajna Soni, Deepika Raman, Dylan Hadfield-Menell
Datasets play a central role in AI governance by enabling both evaluation (measuring capabilities) and alignment (enforcing values) along axes such as helpfulness, harmlessness, to…