1 paper
Eric S. Qiu, Utku Umur Acikalin, Justin Lovelace +4
Mixture-of-Experts (MoE) architectures decouple model capacity from computational cost, yet incur high memory footprints as parameters grow linearly with the number of experts. Rec…