3 papers
cs.DC2026
HyperParallel-Mpipe: A Composable Algebra System for Optimizing MLLM Training over Supernode Clusters
Chong Li, Zhengdao Yu, Nelson Lossing +6
Modern AI applications have expanded beyond text-only interaction into a wide range of multimodal scenarios, making multimodal large language models (MLLMs) crucial for both resear…
cs.DC2026
HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs
Zewen Jin, Congkun Ai, Guangpeng Zhang +7
Modern Mixture-of-Experts (MoE) models increasingly rely on large-scale AI accelerator clusters for efficient training. Ascend NPUs expose heterogeneous on-chip compute resources,…
cs.CL2025
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…