1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2025
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning
Giang Do, Hung Le, Truyen Tran
Sparse Mixture of Experts (SMoE) enables efficient training of large language models by routing input tokens to a select number of experts. However, training SMoE remains challengi…
cs.CL2024★ 1 cited
SimSMoE: Solving Representational Collapse via Similarity Measure
Giang Do, Hung Le, Truyen Tran
Sparse mixture of experts (SMoE) have emerged as an effective approach for scaling large language models while keeping a constant computational cost. Regardless of several notable…
cs.LG2024★ 1 cited
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
Quang Pham, Giang Do, Huy Nguyen +8
Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, effective traini…