4 citations · 6 across the 5 of their papers we have counts for
1 paper · 2 filters
Nayeon Kim, Hojin Lee, Yunju Bak +2
Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---partic…