1 paper
Matan Solomon, Ofra Amir, Omer Ben-Porat
Reinforcement learning agents are often updated with human feedback, yet such updates can be unreliable: reward misspecification, preference conflicts, or limited data may leave po…