1 paper
Aaron Broukhim, Nadir Weibel, Eshin Jolly
Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for…