3 papers
cs.LG2026
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
Tianhao Hu, Xiangcheng Liu, Youshao Xiao +21
Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed ge…
cs.LG2026
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
Zhiyuan Yao, Yi-Kai Zhang, Yuxin Chen +7
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key approach for enhancing LLM reasoning. However, standard frameworks like Group Relative Policy Optimizatio…
cs.CL2025
Self-playing Adversarial Language Game Enhances LLM Reasoning
Pengyu Cheng, Tianhao Hu, Han Xu +6
We explore the potential of self-play training for large language models (LLMs) in a two-player adversarial language game called Adversarial Taboo. In this game, an attacker and a…