3 citations · 6 across the 9 of their papers we have counts for
1 paper · 2 filters
Daniel Fein, Max Lamparth, Violet Xiang +2
Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby L…