Accelerating MoE Model Inference with Expert Sharding
arXiv:2503.08467 · doi:10.1145/3721146.3721938
Abstract
Mixture of experts (MoE) models achieve state-of-the-art results in language modeling but suffer from inefficient hardware utilization due to imbalanced token routing and communication overhead. While prior work has focused on optimizing MoE training and decoder architectures, inference for encoder-based MoE models in a multi-GPU with expert parallelism setting remains underexplored. We introduce MoEShard, an inference system that achieves perfect load balancing through tensor sharding of MoE experts. Unlike existing approaches that rely on heuristic capacity factors or drop tokens, MoEShard evenly distributes computation across GPUs and ensures full token retention, maximizing utilization regardless of routing skewness. We achieve this through a strategic row- and column-wise decomposition of expert matrices. This reduces idle time and avoids bottlenecks caused by imbalanced expert assignments. Furthermore, MoEShard minimizes kernel launches by fusing decomposed expert computations, significantly improving throughput. We evaluate MoEShard against DeepSpeed on encoder-based architectures, demonstrating speedups of up to 6.4 in time to first token (TTFT). Our results show that tensor sharding, when properly applied to experts, is a viable and effective strategy for efficient MoE inference.
To appear in the proceedings of the 5th Workshop on Machine Learning and Systems (EuroMLSys 25)
References in corpus (17)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Towards Federated Learning at Scale: System Design
- Evaluating Gradient Inversion Attacks and Defenses in Federated Learning
- Resource-Efficient Federated Learning
- Oort: Efficient Federated Learning via Guided Participant Selection
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Federated Learning with Buffered Asynchronous Aggregation
- Papaya: Practical, Private, and Scalable Federated Learning
- Beyond spectral gap: The role of the topology in decentralized learning
- Decentralized Learning Made Easy with DecentralizePy
- Decentralized Federated Learning: A Survey and Perspective
- Noiseless Privacy-Preserving Decentralized Learning
- Exponential Graph is Provably Efficient for Decentralized Deep Training
- Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time Convergence
- Communication-Efficient Topologies for Decentralized Learning with Consensus Rate
- PeerSwap: A Peer-Sampler with Randomness Guarantees
- Scalable Decentralized Learning with Teleportation