A Survey on Mixture of Experts in Large Language Models
arXiv:2407.06204 · doi:10.1109/TKDE.2025.3554028
Abstract
Large language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond. The prowess of LLMs is underpinned by their substantial model size, extensive and diverse datasets, and the vast computational power harnessed during training, all of which contribute to the emergent abilities of LLMs (e.g., in-context learning) that are not present in small models. Within this context, the mixture of experts (MoE) has emerged as an effective method for substantially scaling up model capacity with minimal computation overhead, gaining significant attention from academia and industry. Despite its growing prevalence, there lacks a systematic and comprehensive review of the literature on MoE. This survey seeks to bridge that gap, serving as an essential resource for researchers delving into the intricacies of MoE. We first briefly introduce the structure of the MoE layer, followed by proposing a new taxonomy of MoE. Next, we overview the core designs for various MoE models including both algorithmic and systemic aspects, alongside collections of available open-source implementations, hyperparameter configurations and empirical evaluations. Furthermore, we delineate the multifaceted applications of MoE in practice, and outline some potential directions for future research. To facilitate ongoing updates and the sharing of cutting-edge advances in MoE research, we have established a resource repository at https://github.com/withinmiaov/A-Survey-on-Mixture-of-Experts-in-LLMs.
The first three authors contributed equally to this work; Accepted by TKDE
References in corpus (7)
- No Language Left Behind: Scaling Human-Centered Machine Translation
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- AdaMCT: Adaptive Mixture of CNN-Transformer for Sequential Recommendation
- A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
- LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
Cited by in corpus (9)
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Towards Incremental Learning in Large Language Models: A Critical Review
- Mixture of Experts for Decentralized Generative AI and Reinforcement Learning in Wireless Networks: A Comprehensive Survey
- A Glass-Box Deep-Learning Method for Electrical Energy System Modeling Based on Kolmogorov-Arnold Network
- A Deep Generative Model for Five-Class Sleep Staging with Arbitrary Sensor Input
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- MoIRA: Modular Instruction Routing Architecture for Multi-Task Robotics
- A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
- Trust-Aware Routing for Distributed Generative AI Inference at the Edge