10 citations · 20 across the 3 of their papers we have counts for
1 paper · 1 filter
Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4
Large language models are typically aligned with human preferences by optimizing reward models (RMs) fitted to human feedback. However, human preferences are multi-facet…