4 citations · 10 across the 29 of their papers we have counts for
8 papers · 2 filters
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
Pedram Akbarian, Huy Nguyen, Xing Han +1
Mixture of Experts (MoE) models are well known for effectively scaling model capacity while preserving computational overheads. In this paper, we establish a rigorous relation betw…
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
Huy Nguyen, Xing Han, Carl Harris +2
With the growing prominence of the Mixture of Experts (MoE) architecture in developing large-scale foundation models, we investigate the Hierarchical Mixture of Experts (HMoE), a s…
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
Huy Nguyen, Pedram Akbarian, Trang Pham +3
The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable…
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
Huy Nguyen, Nhat Ho, Alessandro Rinaldo
The softmax gating function is arguably the most popular choice in mixture of experts modeling. Despite its widespread use in practice, the softmax gating may lead to unnecessary c…
On Parameter Estimation in Deviated Gaussian Mixture of Experts
Huy Nguyen, Khai Nguyen, Nhat Ho
We consider the parameter estimation problem in the deviated Gaussian mixture of experts in which the data are generated from $(1 - λ^{\ast}) g_0(Y| X)+ λ^{\ast} \sum_{i = 1}^{k_{\…
On Least Square Estimation in Softmax Gating Mixture of Experts
Huy Nguyen, Nhat Ho, Alessandro Rinaldo
Mixture of experts (MoE) model is a statistical machine learning design that aggregates multiple expert networks using a softmax gating function in order to form a more intricate a…