1 paper
Gyeongman Kim, Gyouk Chu, Eunho Yang
With the emergence of Mixture-of-Experts (MoE), the efficient scaling of model size has accelerated the development of large language models in recent years. However, their high me…