16 citations · 16 across the 2 of their papers we have counts for
1 paper · 1 filter
Nicholas Stranges, Yimin Yang
Current Large Language Model (LLM) preference learning methods such as Proximal Policy Optimization and Direct Preference Optimization learn from direct rankings or numerical ratin…