1 paper · 1 filter
Zhanhui Zhou, Jie Liu, Jing Shao +4
A single language model, even when aligned with labelers through reinforcement learning from human feedback (RLHF), may not suit all human preferences. Recent approaches therefore…