2 papers
cs.DC2026
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
Minyu Cui, Anna Wingkvist, Morgan Ericsson
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language mo…
cs.DC2026
Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads
Minyu Cui, Miquel Pericas
The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems. As model sizes and computatio…