9 citations · 15 across the 31 of their papers we have counts for
1 paper · 2 filters
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2
To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effecti…