Advances in Preference-based Reinforcement Learning: A Review
arXiv:2408.11943 · doi:10.1109/SMC53654.2022.9945333
Abstract
Reinforcement Learning (RL) algorithms suffer from the dependency on accurately engineered reward functions to properly guide the learning agents to do the required tasks. Preference-based reinforcement learning (PbRL) addresses that by utilizing human preferences as feedback from the experts instead of numeric rewards. Due to its promising advantage over traditional RL, PbRL has gained more focus in recent years with many significant advances. In this survey, we present a unified PbRL framework to include the newly emerging approaches that improve the scalability and efficiency of PbRL. In addition, we give a detailed overview of the theoretical guarantees and benchmarking work done in the field, while presenting its recent applications in complex real-world tasks. Lastly, we go over the limitations of the current approaches and the proposed future research directions.
References in corpus (11)
- A Simple Framework for Contrastive Learning of Visual Representations
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Trust Region Policy Optimization
- DeepMind Control Suite
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- A Survey of Deep Reinforcement Learning in Video Games
- Recursively Summarizing Books with Human Feedback
- Improved Optimistic Algorithms for Logistic Bandits
- Maximum Selection and Ranking under Noisy Comparisons
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
- Preference-based Reinforcement Learning with Finite-Time Guarantees