1 paper
Jiaxin Wen, Ruiqi Zhong, Akbir Khan +6
Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this p…