9 papers
Normalized Rewards for Preference Optimization
Shawn Im, Federico Danieli, Skyler Seto +2
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their…
ExpertLens: Activation steering features are highly interpretable
Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson +3
Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amoun…
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
Anirudh Sundar, Sinead Williamson, Katherine Metcalf +3
Aligned representations across languages is a desired property in multilingual large language models (mLLMs), as alignment can improve performance in cross-lingual tasks. Typically…
Aligning LLMs by Predicting Preferences from User Writing Samples
Stéphane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald +1
Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acti…
On the Way to LLM Personalization: Learning to Remember User Conversations
Lucie Charlotte Magister, Katherine Metcalf, Yizhe Zhang +1
Large Language Models (LLMs) have quickly become an invaluable assistant for a variety of tasks. However, their effectiveness is constrained by their ability to tailor responses to…
PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
Stephane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald +1
Accommodating human preferences is essential for creating AI agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs to infer pref…