5 papers
A Dataset for Dynamic Human Preferences for Vision Language Models
Hannah Gao, Dylan Hadfield-Menell, Rachel Ma
Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models can adapt to real-time preferenc…
A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers
Hannah Gao, Isha Agarwal, Dylan Hadfield-Menell +1
The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances or perturbations, such as blurr…
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
Rachel Ma, Dylan Hadfield-Menell, Kristjan Greenewald
Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the fir…
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
Rachel Ma, Jingyi Qu, Andreea Bobu +1
We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferen…
Goal Inference from Open-Ended Dialog
Rachel Ma, Jingyi Qu, Andreea Bobu +1
Embodied AI Agents are quickly becoming important and common tools in society. These embodied agents should be able to learn about and accomplish a wide range of user goals and pre…