Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
Pedram Akbarian, Huy Nguyen, Xing Han +1
Mixture of Experts (MoE) models are well known for effectively scaling model capacity while preserving computational overheads. In this paper, we establish a rigorous relation betw…
stat.ML2025
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
Huy Nguyen, Xing Han, Carl Harris +2
With the growing prominence of the Mixture of Experts (MoE) architecture in developing large-scale foundation models, we investigate the Hierarchical Mixture of Experts (HMoE), a s…