2 papers
cs.LG2026
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
Mohamed Nabail, Leo Kaixuan Cheng, Jingmin Wang +1
Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design. However, existing methods…
cs.LG2025
Residual Reward Models for Preference-based Reinforcement Learning
Chenyang Cao, Miguel Rogel-GarcÃa, Mohamed Nabail +2
Preference-based Reinforcement Learning (PbRL) provides a way to learn high-performance policies in environments where the reward signal is hard to specify, avoiding heuristic and…