3 papers
cs.LG2026
PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint
Haotian Xie, Junlin Chen, Mingkai Zheng +2
State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware…
cs.LG2026
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Daize Dong, Junlin Chen, Haolong Jia +9
Mixture of Experts (MoE) Large Language Models (LLMs) achieve strong performance at scale. However, reinforcement learning (RL) on MoE-based LLMs often suffers from training instab…
cs.DC2026
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
Zizhao Mo, Junlin Chen, Huanle Xu +1
Nowadays, service providers often deploy multiple types of LLM services within shared clusters. While the service colocation improves resource utilization, it introduces significan…