1 citations · 1 across the 5 of their papers we have counts for
1 paper · 2 filters
Giang Do, Khiem Le, Quang Pham +7
By routing input tokens to only a few split experts, Sparse Mixture-of-Experts has enabled efficient training of large language models. Recent findings suggest that fixing the rout…