1 paper
Brenden Latham, Mehrdad Moharrami
We study offline constrained reinforcement learning from human feedback with multiple preference oracles. Motivated by applications that trade off performance with safety or fairne…