1 paper · 1 filter
Vignesh Adhinarayanan, Nuwan Jayasena
Mixture-of-Experts (MoE) models deliver high quality at low training FLOPs, but this efficiency often vanishes at inference. We identify a double penalty that structurally disadvan…