1 paper
Byeongju Kim, Jungwan Lee, Donghyeon Han +2
Recently, Mixture-of-Experts (MoE) models have gained attention for efficiently scaling large language models. Although these models are extremely large, their sparse activation en…