3 papers
cs.DC2026
Belayer: Efficient Fault Tolerance for LLM Agentic RL Training
Jiecheng Zhou, Qinghao Hu, Peng Sun +2
Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-inten…
cs.AI2025
RL in the Wild: Characterizing RLVR Training in LLM Deployment
Jiecheng Zhou, Qinghao Hu, Yuyang Jin +7
Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent month…
cs.DC2025
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
Chang Chen, Tiancheng Chen, Jiangfei Duan +7
Training large language models (LLMs) with increasingly long and varying sequence lengths introduces severe load imbalance challenges in large-scale data-parallel training. Recent…