1 paper
Zheng Zhang, Donglin Yang, Yaqi Xia +4
Recently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows fo…