Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
Chenglong Wang, Ziming Zhu, Yifu Huo +9
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranki…
cs.LG2026
FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
Kaiyang Ye, Yuan Ge, Junxiang Zhang +8
While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored.…
cs.LG2025
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
Yongyu Mu, Jiali Zeng, Bei Li +5
Despite recent progress in training long-chain-of-thought reasoning models via scaling reinforcement learning (RL), its underlying training dynamics remain poorly understood, and s…