1 paper · 1 filter
Rongfeng Wang, Shichao Weng, Zhiqiang Wang +4
Mixture-of-Experts (MoE) models route each token to a subset of expert networks, increasing capacity while keeping per-token computation sparse. In many deployed MoEs, the number o…