1 paper · 1 filter
Yu Gong, Kailash Budhathoki, Taeho Kim +2
Mixture-of-Experts (MoE) layers increase model capacity without proportionally increasing arithmetic, but their sparse expert computation is difficult to execute efficiently during…