6 papers
Co-Evolving LLM Evaluators and Policies via DynamicRubric
Beining Wang, Weihang Su, Hongtao Tian +8
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become…
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners
Chao Wang, Hongtao Tian, Tao Yang +3
Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense process supervision -- via learned process r…
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
Guofu Xie, Yunsheng Shi, Hongtao Tian +2
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of Large Language Models (LLMs) by using rule-based binary feedback. However, current RLV…
From Faithfulness to Correctness: Generative Reward Models that Think Critically
Qiyao Ma, Yunsheng Shi, Hongtao Tian +3
Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathemat…
WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
Junyu Wu, Weiming Chang, Xiaotao Liu +10
Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent paradigm for training large language models and multimodal systems. Despite the notable advances enable…
G-Core: A Simple, Scalable and Balanced RLHF Trainer
Junyu Wu, Weiming Chang, Xiaotao Liu +8
Reinforcement Learning from Human Feedback (RLHF) has become an increasingly popular paradigm for training large language models (LLMs) and diffusion models. While existing RLHF tr…