1 paper
Udita Ghosh, Dripta S. Raychaudhuri, Jiachen Li +2
Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In thi…