1 paper
Joongwon Kim, Anirudh Goyal, Aston Zhang +6
Preference learning is a widely adopted post-training technique that aligns large language models (LLMs) to human preferences and improves specific downstream task capabilities. In…