activity
20242026
collaborators

15 papers

cs.LG2026

Normalized Rewards for Preference Optimization

Shawn Im, Federico Danieli, Skyler Seto +2

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their…

cs.LG2026

Theoretical Limits of Language Model Alignment

Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1

Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i)…

cs.LG2026

DSO: Direct Steering Optimization for Bias Mitigation

Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang +4

Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually imp…

cs.CL2025

Bias after Prompting: Persistent Discrimination in Large Language Models

Nivedha Sivakumar, Natalie Mackraz, Samira Khorshidi +4

A dangerous assumption that can be made from prior work on the bias transfer hypothesis (BTH) is that biases do not transfer from pre-trained large language models (LLMs) to adapte…

cs.CL2025

ExpertLens: Activation steering features are highly interpretable

Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson +3

Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amoun…

cs.CL2025

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang +5

Large language models (LLMs) have achieved impressive performance, leading to their widespread adoption as decision-support tools in resource-constrained contexts like hiring and a…