18 citations · 23 across the 7 of their papers we have counts for
1 paper · 1 filter
Giang Do, Hung Le, Truyen Tran
Sparse mixture of experts (SMoE) have emerged as an effective approach for scaling large language models while keeping a constant computational cost. Regardless of several notable…