1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Yannick Metz, András Geiszl, Raphaël Baur +1
Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison betwe…