1 paper
Giang Do, Khiem Le, Quang Pham +7
By routing input tokens to only a few split experts, Sparse Mixture-of-Experts has enabled efficient training of large language models. Recent findings suggest that fixing the rout…