1 paper · 1 filter
Ethan He, Abhinav Khattar, Ryan Prenger +7
Upcycling pre-trained dense language models into sparse mixture-of-experts (MoE) models is an efficient approach to increase the model capacity of already trained models. However,…