1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Adrian Zhao, Zhenkun Cai, Zhenyu Song +5
Mixture-of-Experts (MoE) has recently emerged as the mainstream architecture for efficiently scaling large language models while maintaining near-constant computational cost. Exper…