2 citations · 2 across the 1 of their papers we have counts for
1 paper
Maryam Akhavan Aghdam, Hongpeng Jin, Yanzhao Wu
Transformer-based Mixture-of-Experts (MoE) models have been driving several recent technological advancements in Natural Language Processing (NLP). These MoE models adopt a router…