3 papers
cs.LG2025
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
Chao Wang, Tao Yang, Hongtao Tian +5
Critic-free methods like GRPO reduce memory demands by estimating advantages from multiple rollouts but tend to converge slowly, as critical learning signals are diluted by an abun…
cs.LG2025
WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
Junyu Wu, Weiming Chang, Xiaotao Liu +10
Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent paradigm for training large language models and multimodal systems. Despite the notable advances enable…
cs.LG2025
G-Core: A Simple, Scalable and Balanced RLHF Trainer
Junyu Wu, Weiming Chang, Xiaotao Liu +8
Reinforcement Learning from Human Feedback (RLHF) has become an increasingly popular paradigm for training large language models (LLMs) and diffusion models. While existing RLHF tr…