3 papers
cs.LG2026
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
Tianhao Hu, Xiangcheng Liu, Youshao Xiao +21
Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed ge…
cs.LG2025
Decoding Large Language Diffusion Models with Foreseeing Movement
Yichuan Mo, Quan Chen, Mingjie Li +2
Large Language Diffusion Models (LLDMs) benefit from a flexible decoding mechanism that enables parallelized inference and controllable generations over autoregressive models. Yet…
cs.CL2025
Advancing LLM Safe Alignment with Safety Representation Ranking
Tianqi Du, Zeming Wei, Quan Chen +2
The rapid advancement of large language models (LLMs) has demonstrated milestone success in a variety of tasks, yet their potential for generating harmful content has raised signif…