1 paper · 1 filter
Xiandong Zou, Wanyu Lin, Yuchen Li +1
Aligning Large Language Model (LLM) responses with human preferences is vital for building safe and controllable AI systems. While preference optimization methods based on Plackett…