1 paper
Giang Do, Hung Le, Truyen Tran
Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected subset of specialized experts.…