Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Eigenvectors of Experts are Training-free Non-collapsing Routers
Giang Do, Hung Le, Truyen Tran
Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected subset of specialized experts.…
cs.LG2025
On the Role of Discrete Representation in Sparse Mixture of Experts
Giang Do, Kha Pham, Hung Le +1
Sparse mixture of experts (SMoE) is an effective solution for scaling up model capacity without increasing the computational costs. A crucial component of SMoE is the router, respo…