1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Yongxiang Lyu, Ning Li, Bonian Jia
Mixture-of-Experts (MoE) models are increasingly used in LLMs because sparse activation decouples model capacity from compute cost. However, the large expert parameter footprint of…