25 citations · 48 across the 12 of their papers we have counts for
1 paper · 1 filter
Hao Sun, Yunyi Shen, Jean-Francois Ton
The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear why this model -- original…