1 paper
Teng Xiao, Yige Yuan, Zhengyu Chen +4
Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasin…