1 paper
Osayamen Jonathan Aimuyo, Byungsoo Oh, Rachee Singh
The computational sparsity of Mixture-of-Experts (MoE) models enables sub-linear growth in compute cost as model size increases, thus offering a scalable path to training massive n…