30 citations · 96 across the 36 of their papers we have counts for
1 paper · 1 filter
Prasann Singhal, Nathan Lambert, Scott Niekum +2
Varied approaches for aligning language models have been proposed, including supervised fine-tuning, RLHF, and direct optimization methods such as DPO. Although DPO has rapidly gai…