3 citations · 3 across the 8 of their papers we have counts for
1 paper · 2 filters
Yifan Zhong, Chengdong Ma, Xiaoyuan Zhang +5
Current methods for large language model alignment typically use scalar human preference labels. However, this convention tends to oversimplify the multi-dimensional and heterogene…