1 paper · 1 filter
Hao Sun, Yunyi Shen, Jean-Francois Ton
The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear why this model -- original…