430 citations · 558 across the 8 of their papers we have counts for
1 paper · 1 filter
Kenan Hasanaliyev, Silas Alberti, Jenny Hamer +5
Reward models are a standard tool to score responses from LLMs. Reward models are built to rank responses to a fixed prompt sampled from a single model, for example to choose the b…