6 citations · 8 across the 31 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Jiazheng Zhang, Wenqing Jing, Zizhuo Zhang +9
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human values. However, noisy preferences in human feedback can lead to reward misgeneralizatio…
cs.LG2025
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
Deyang Kong, Qi Guo, Xiangyu Xi +5
Reinforcement learning exhibits potential in enhancing the reasoning abilities of large language models, yet it is hard to scale for the low sample efficiency during the rollout ph…
cs.LG2024
Length Desensitization in Direct Preference Optimization
Wei Liu, Yang Bai, Chengcheng Han +5
Direct Preference Optimization (DPO) is widely utilized in the Reinforcement Learning from Human Feedback (RLHF) phase to align Large Language Models (LLMs) with human preferences,…