15 papers
Normalized Rewards for Preference Optimization
Shawn Im, Federico Danieli, Skyler Seto +2
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their…
Theoretical Limits of Language Model Alignment
Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1
Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i)…
DSO: Direct Steering Optimization for Bias Mitigation
Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang +4
Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually imp…
Bias after Prompting: Persistent Discrimination in Large Language Models
Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4
A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…
ExpertLens: Activation steering features are highly interpretable
Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson +3
Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amoun…
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang +5
Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and a…