1 paper
Shwai He, Run-Ze Fan, Liang Ding +3
Scaling the size of language models usually leads to remarkable advancements in NLP tasks. But it often comes with a price of growing computational cost. Although a sparse Mixture…