Deep reinforcement learning from human preferences
arXiv:1706.03741
Abstract
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effectively solve complex RL tasks without access to the reward function, including Atari games and simulated robot locomotion, while providing feedback on less than one percent of our agent's interactions with the environment. This reduces the cost of human oversight far enough that it can be practically applied to state-of-the-art RL systems. To demonstrate the flexibility of our approach, we show that we can successfully train complex novel behaviors with about an hour of human time. These behaviors and environments are considerably more complex than any that have been previously learned from human feedback.
Cited by in corpus (88)
- An Introduction to Deep Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Thinking Fast and Slow in Large Language Models
- Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4
- Decoding ChatGPT: A Taxonomy of Existing Research, Current Challenges, and Possible Future Directions
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- Factuality Challenges in the Era of Large Language Models
- A Survey of Deep Reinforcement Learning in Video Games
- Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review
- Machine Culture
- Scalable agent alignment via reward modeling: a research direction
- AI Safety Gridworlds
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
- Aligning AI With Shared Human Values
- Learning from models beyond fine-tuning
- GPT-4 can pass the Korean National Licensing Examination for Korean Medicine Doctors
- DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback
- Reward learning from human preferences and demonstrations in Atari
- Few-Shot Goal Inference for Visuomotor Learning and Planning
- AI Safety for Everyone
- Deep Reinforcement Learning from Policy-Dependent Human Feedback
- Beyond Preferences in AI Alignment
- Designing Deep Reinforcement Learning for Human Parameter Exploration
- Explore, Exploit or Listen: Combining Human Feedback and Policy Model to Speed up Deep Reinforcement Learning in 3D Worlds
- A General Language Assistant as a Laboratory for Alignment
- Supervising strong learners by amplifying weak experts
- AI Research Considerations for Human Existential Safety (ARCHES)
- Hierarchical Imitation and Reinforcement Learning
- Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
- Batch Active Preference-Based Learning of Reward Functions
- Inverse reinforcement learning for video games
- Risk-Aware Active Inverse Reinforcement Learning
- Model Inversion Networks for Model-Based Optimization
- Standards for Belief Representations in LLMs
- Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function)
- Conservative Agency via Attainable Utility Preservation
- Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
- Cycle-of-Learning for Autonomous Systems from Human Interaction
- FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
- Discovering Blind Spots in Reinforcement Learning
- Scalable Psychological Momentum Forecasting in Esports
- Learning Norms from Stories: A Prior for Value Aligned Agents
- Truthful AI: Developing and governing AI that does not lie
- Requisite Variety in Ethical Utility Functions for AI Value Alignment
- Learning via social awareness: Improving a deep generative sketching model with facial feedback
- Meta-learners' learning dynamics are unlike learners'
- Guiding Policies with Language via Meta-Learning
- Normative Conflicts and Shallow AI Alignment
- Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback
- Using Monte Carlo Tree Search as a Demonstrator within Asynchronous Deep RL
- Stick to your Role! Stability of Personal Values Expressed in Large Language Models
- Active Reinforcement Learning with Monte-Carlo Tree Search
- Hidden Incentives for Auto-Induced Distributional Shift
- Just Ask:An Interactive Learning Framework for Vision and Language Navigation
- APRIL: Interactively Learning to Summarise by Combining Active Preference Learning and Reinforcement Learning
- Active Inverse Reward Design
- Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
- DeepCrawl: Deep Reinforcement Learning for Turn-based Strategy Games
- Computational Rational Engineering and Development: Synergies and Opportunities
- 'Indifference' methods for managing agent rewards
- Safe Driving via Expert Guided Policy Optimization
- Understanding Learned Reward Functions
- Choice Set Misspecification in Reward Inference
- Leveraging Human Guidance for Deep Reinforcement Learning Tasks
- Parenting: Safe Reinforcement Learning from Human Input
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- Avoiding Side Effects in Complex Environments
- Improving image generative models with human interactions
- EvoText: Enhancing Natural Language Generation Models via Self-Escalation Learning for Up-to-Date Knowledge and Improved Performance
- Differential-Critic GAN: Generating What You Want by a Cue of Preferences
- Offline Preference-Based Apprenticeship Learning
- Policy Gradient from Demonstration and Curiosity
- Parallelized Interactive Machine Learning on Autonomous Vehicles
- Transfer Reward Learning for Policy Gradient-Based Text Generation
- Deep Reinforcement Learning for Playing 2.5D Fighting Games
- Inverse Reinforcement Learning via Matching of Optimality Profiles
- Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
- Region Growing Curriculum Generation for Reinforcement Learning
- Variational Policy Search using Sparse Gaussian Process Priors for Learning Multimodal Optimal Actions
- Assisted Robust Reward Design
- Teachable Reinforcement Learning via Advice Distillation
- RL agents Implicitly Learning Human Preferences
- Cost Functions for Robot Motion Style
- Robby is Not a Robber (anymore): On the Use of Institutions for Learning Normative Behavior
- Enhancing AI Safety Through the Fusion of Low Rank Adapters