2 citations · 2 across the 1 of their papers we have counts for
1 paper
Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati
Preference Based Reinforcement Learning has shown much promise for utilizing human binary feedback on queried trajectory pairs to recover the underlying reward model of the Human i…