1 paper · 1 filter
Xubo Liu, Wenya Guo, Ruxue Yan +2
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for thi…