1 paper
Zhiwei He, Xing Wang, Wenxiang Jiao +4
Insufficient modeling of human preferences within the reward model is a major obstacle for leveraging human feedback to improve translation quality. Fortunately, quality estimation…