Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models
Yongqin Zeng, Sicheng Pan, Jiale Wang +4
Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data rem…
cs.AI2026
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
MiniMax, :, Aili Chen +219
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…