2 papers
cs.CL2025
Expert-Token Resonance MoE: Bidirectional Routing with Efficiency Affinity-Driven Active Selection
Jing Li, Zhijie Sun, Dachao Lin +5
Mixture-of-Experts (MoE) architectures enable efficient scaling of large language models by activating only a subset of parameters per input. However, existing MoE models suffer fr…
cs.LG2024
LocMoE: A Low-Overhead MoE for Large Language Model Training
Jing Li, Zhijie Sun, Xuan He +6
The Mixtures-of-Experts (MoE) model is a widespread distributed and integrated learning method for large language models (LLM), which is favored due to its ability to sparsify and…