4 citations · 4 across the 6 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
Hang Ding, Qiming Feng, Dongqi Liu +9
Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play,…
cs.CL2025
Self-supervised Attribute-aware Dynamic Preference Ranking Alignment
Hongyu Yang, Qi Zhao, Zhenhua hu +1
Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely…