activity
20212026
most citedOpen Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

95 citations · 107 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

Rachel Freedman

Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time. We introduce Adaptive Pluralistic Alignm…

cs.LG2024

Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Vincent Conitzer, Rachel Freedman, Jobst Heitzig +9

Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tu…

cs.AI2023

Active teacher selection for reward learning

Rachel Freedman, Justin Svegliato, Kyle Wray +1

Reward learning techniques enable machine learning systems to learn objectives from human feedback. A core limitation of these systems is their assumption that all feedback comes f…

cs.AI2023★ 95 cited

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Stephen Casper, Xander Davies, Claudia Shi +29

Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…

cs.LG2023★ 2 cited

Active Reward Learning from Multiple Teachers

Peter Barnett, Rachel Freedman, Justin Svegliato +1

Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system. This human feedback is often a preference comparison, in whi…

cs.LG2022★ 6 cited

The Expertise Problem: Learning from Specialized Feedback

Oliver Daniels-Koch, Rachel Freedman

Reinforcement learning from human feedback (RLHF) is a powerful technique for training agents to perform difficult-to-specify tasks. However, human feedback can be noisy, particula…