Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Eigenvectors of Experts are Training-free Non-collapsing Routers
Giang Do, Hung Le, Truyen Tran
Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected subset of specialized experts.…
cs.CL2026
Do Domain-specific Experts exist in MoE-based LLMs?
Giang Do, Hung Le, Truyen Tran
In the era of Large Language Models (LLMs), the Mixture of Experts (MoE) architecture has emerged as an effective approach for training extremely large models with improved computa…