7 papers
A Geometric Perspective on Stabilizing Value Conflict Resolution
Saket Reddy, Andy Liu
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To add…
Generative Value Conflicts Reveal LLM Priorities
Andy Liu, Kshitish Ghate, Mona Diab +3
Past work seeks to align large language model (LLM)-based assistants with a target set of values, but such assistants are frequently forced to make tradeoffs between values when de…
Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy
Wenkai Li, Lynnette Hui Xian Ng, Andy Liu +1
The study of negotiation styles dates back to Aristotle's ethos-pathos-logos rhetoric. Prior efforts primarily studied the success of negotiation agents. Here, we shift the focus t…
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
Kshitish Ghate, Andy Liu, Devansh Jain +5
As large language models (LLMs) are deployed globally, creating pluralistic systems that can accommodate the diverse preferences and values of users worldwide becomes essential. We…
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
Wenkai Li, Jiarui Liu, Andy Liu +3
In this work, we tackle the challenge of embedding realistic human personality traits into LLMs. Previous approaches have primarily focused on prompt-based methods that describe th…
Evaluating LLM Agent Collusion in Double Auctions
Kushal Agrawal, Verona Teo, Juan J. Vazquez +3
Large language models (LLMs) have demonstrated impressive capabilities as autonomous agents with rapidly expanding applications in various domains. As these agents increasingly eng…