3 papers
cs.CL2025
A Survey on LLM Mid-Training
Chengying Tu, Xuemiao Zhang, Rongxiang Weng +6
Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage…
cs.CL2025
Libra: Assessing and Improving Reward Model by Learning to Think
Meng Zhou, Bei Li, Jiahao Liu +5
Reinforcement learning (RL) has significantly improved the reasoning ability of large language models. However, current reward models underperform in challenging reasoning scenario…
cs.LG2024
Length Desensitization in Direct Preference Optimization
Wei Liu, Yang Bai, Chengcheng Han +5
Direct Preference Optimization (DPO) is widely utilized in the Reinforcement Learning from Human Feedback (RLHF) phase to align Large Language Models (LLMs) with human preferences,…