1 paper
Giang Do, Hung Le, Truyen Tran
Sparse mixture of experts (SMoE) have emerged as an effective approach for scaling large language models while keeping a constant computational cost. Regardless of several notable…