1 paper · 1 filter
Shaoyu Wang, Guangrong He, Geon-Woo Kim +2
Mixture-of-Experts (MoE) architectures offer the promise of larger model capacity without the prohibitive costs of fully dense designs. However, in real-world inference serving, lo…