Showing cs.DCShow all
3 papers · 1 filter
cs.DC2026
Belayer: Efficient Fault Tolerance for LLM Agentic RL Training
Jiecheng Zhou, Qinghao Hu, Peng Sun +2
Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-inten…
cs.DC2025
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
Chang Chen, Tiancheng Chen, Jiangfei Duan +7
Training large language models (LLMs) with increasingly long and varying sequence lengths introduces severe load imbalance challenges in large-scale data-parallel training. Recent…
cs.DC2024★ 6 cited
Characterization of Large Language Model Development in the Datacenter
Qinghao Hu, Zhisheng Ye, Zerui Wang +9
Large Language Models (LLMs) have presented impressive performance across several transformative tasks. However, it is non-trivial to efficiently utilize large-scale cluster resour…