1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Hang Ding, Qiming Feng, Dongqi Liu +9
Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play,…