1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Yannick Metz, David Lindner, Raphaël Baur +2
To use reinforcement learning from human feedback (RLHF) in practical applications, it is crucial to learn reward models from diverse sources of human feedback and to consider huma…