2 citations · 2 across the 1 of their papers we have counts for
1 paper
Pengyu Cheng, Jiawen Xie, Ke Bai +2
Reward models (RMs) are essential for aligning large language models (LLMs) with human preferences to improve interaction quality. However, the real world is pluralistic, which lea…