2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Yongji Wu, Xueshen Liu, Shuowei Jin +6
The Mixture-of-Experts (MoE) architecture has become increasingly popular as a method to scale up large language models (LLMs). To save costs, heterogeneity-aware training solution…