4 papers
Analytical Resource Management for Fine-grained MoE Computation-Communication Overlap
Hongyu Liu, Minyu Cui, Miquel Pericas
Fine-grained computation--communication overlap in distributed Mixture-of-Experts (MoE) inference allows communication to begin as partial compute results become ready. However, co…
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
Minyu Cui, Anna Wingkvist, Morgan Ericsson
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language mo…
Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads
Minyu Cui, Miquel Pericas
The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems. As model sizes and computatio…
Analysis and Characterization of Performance Variability for OpenMP Runtime
Minyu Cui, Nikela Papadopoulou, Miquel Pericàs
In the high performance computing (HPC) domain, performance variability is a major scalability issue for parallel computing applications with heavy synchronization and communicatio…