2 papers
cs.LG2025
Pairwise Calibrated Rewards for Pluralistic Alignment
Daniel Halpern, Evi Micha, Ariel D. Procaccia +1
Current alignment pipelines presume a single, universal notion of desirable behavior. However, human preferences often diverge across users, contexts, and cultures. As a result, di…
cs.GT2024
Axioms for AI Alignment from Human Feedback
Luise Ge, Daniel Halpern, Evi Micha +4
In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on…