1 paper · 1 filter
Teng Xiao, Zhen Ge, Sujay Sanghavi +5
We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in…