5 papers
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Efficient representations for team and imperfect-recall equilibrium computation
Luca Carminati, Brian Hu Zhang, Federico Cacciamani +4
Equilibrium finding in two-player zero-sum games with perfect recall is a well-studied topic that has led to many breakthroughs in computational game theory. This paper aims to gen…
Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
Dena Mujtaba, Brian Hu, Anthony Hoogs +1
The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments.…
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
Jadie Adams, Brian Hu, Emily Veenhuis +5
Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can on…
ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
Bharadwaj Ravichandran, David Joy, Paul Elliott +6
Large language models (LLMs) are increasingly being used as decision aids. However, users have diverse values and preferences that can affect their decision-making, which requires…