1 citations · 1 across the 11 of their papers we have counts for
6 papers · 1 filter
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training
Zehao Chen, Gongxun Li, Tianxiang Ai +9
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The sa…
Policy Improvement Reinforcement Learning
Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10
Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
Xiaodong Lu, Xiaohan Wang, Jiajun Chai +7
Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language models. However, existing RLVR methods uti…
Heterogeneous Agent Collaborative Reinforcement Learning
Zhixia Zhang, Zixuan Huang, Gongxun Li +10
We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies…
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
Zixuan Huang, Yikun Ban, Lean Fu +4
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly depen…
LLMBoost: Make Large Language Models Stronger with Boosting
Zehao Chen, Tianxiang Ai, Yifei Li +11
Ensemble learning of LLMs has emerged as a promising alternative to enhance performance, but existing approaches typically treat models as black boxes, combining the inputs or fina…