activity
20242026
collaborators

9 papers

cs.LG2026

Normalized Rewards for Preference Optimization

Shawn Im, Federico Danieli, Skyler Seto +2

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their…

cs.CL2025

ExpertLens: Activation steering features are highly interpretable

Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson +3

Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amoun…

cs.CL2025

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf +3

Aligned representations across languages is a desired property in multilingual large language models (mLLMs), as alignment can improve performance in cross-lingual tasks. Typically…

cs.CL2025

Aligning LLMs by Predicting Preferences from User Writing Samples

Stéphane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald +1

Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acti…

cs.CL2024

On the Way to LLM Personalization: Learning to Remember User Conversations

Lucie Charlotte Magister, Katherine Metcalf, Yizhe Zhang +1

Large Language Models (LLMs) have quickly become an invaluable assistant for a variety of tasks. However, their effectiveness is constrained by their ability to tailor responses to…

cs.AI2024

PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories

Stephane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald +1

Accommodating human preferences is essential for creating AI agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs to infer pref…