16 citations · 22 across the 5 of their papers we have counts for
1 paper · 1 filter
Nathan Lambert, Valentina Pyatkin, Jacob Morrison +9
Reward models (RMs) are at the crux of successfully using RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluatio…