Showing cs.DCShow all
3 papers · 1 filter
cs.DC2026
HyperParallel-Mpipe: A Composable Algebra System for Optimizing MLLM Training over Supernode Clusters
Chong Li, Zhengdao Yu, Nelson Lossing +6
Modern AI applications have expanded beyond text-only interaction into a wide range of multimodal scenarios, making multimodal large language models (MLLMs) crucial for both resear…
cs.DC2026
HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs
Zewen Jin, Congkun Ai, Guangpeng Zhang +7
Modern Mixture-of-Experts (MoE) models increasingly rely on large-scale AI accelerator clusters for efficient training. Ascend NPUs expose heterogeneous on-chip compute resources,…
cs.DC2024
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
Ao Sun, Weilin Zhao, Xu Han +4
Effective attention modules have played a crucial role in the success of Transformer-based large language models (LLMs), but the quadratic time and memory complexities of these att…