1 paper
Jing Dong, Yaoliang Yu, Pascal Pourpart
In RLHF, each training example contains a prompt x and two candidate responses y,y′, and annotators provide pairwise preferences between these responses. The learning problem i…