1 paper
Pingzhi Li, Xiaolong Jin, Zhen Tan +2
Mixture-of-Experts (MoE) is a promising way to scale up the learning capacity of large language models. It increases the number of parameters while keeping FLOPs nearly constant du…