1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yannick Metz, David Lindner, Raphaël Baur +2
To use reinforcement learning from human feedback (RLHF) in practical applications, it is crucial to learn reward models from diverse sources of human feedback and to consider huma…