95 citations · 107 across the 6 of their papers we have counts for
7 papers
Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy
Rachel Freedman
Prevailing alignment methods target a fixed set of preferences and therefore risk forcing value lock-in as societal norms evolve over time. We introduce Adaptive Pluralistic Alignm…
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig +9
Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tu…
Active teacher selection for reward learning
Rachel Freedman, Justin Svegliato, Kyle Wray +1
Reward learning techniques enable machine learning systems to learn objectives from human feedback. A core limitation of these systems is their assumption that all feedback comes f…
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi +29
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals. RLHF has emerged as the central method used to finetune state-of…
Active Reward Learning from Multiple Teachers
Peter Barnett, Rachel Freedman, Justin Svegliato +1
Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system. This human feedback is often a preference comparison, in whi…
The Expertise Problem: Learning from Specialized Feedback
Oliver Daniels-Koch, Rachel Freedman
Reinforcement learning from human feedback (RLHF) is a powerful technique for training agents to perform difficult-to-specify tasks. However, human feedback can be noisy, particula…