1 paper
Zixu Shen, Kexin Chu, Yifan Zhang +3
The expansion of large language models is increasingly limited by the constrained memory capacity of modern GPUs. To mitigate this, Mixture-of-Experts (MoE) architectures activate…