2 citations · 3 across the 2 of their papers we have counts for
1 paper · 2 filters
Chufan Shi, Cheng Yang, Xinyu Zhu +6
Mixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activat…