1 paper
Shaolong Chen, Madalina Ciobanu, Qingqing Mao +1
DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loop. Recent theoretical analys…