1 paper
Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi +5
The parameter size of modern large language models (LLMs) can be scaled up via the sparsely-activated Mixture-of-Experts (MoE) technique to avoid excessive increase of the computat…