collaborators

7 papers

stat.ML2025

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

Fanqi Yan, Huy Nguyen, Dung Le +3

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstr…

stat.ML2025

Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention

Pedram Akbarian, Huy Nguyen, Xing Han +1

Mixture of Experts (MoE) models are well known for effectively scaling model capacity while preserving computational overheads. In this paper, we establish a rigorous relation betw…

cs.LG2025

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective

Fanqi Yan, Huy Nguyen, Pedram Akbarian +2

At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on th…

cs.LG2025

Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts

Fanqi Yan, Huy Nguyen, Dung Le +2

We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts. This model is motivated from the prompt learning problem where ones utilize prom…

stat.ML2025

Statistical Advantages of Perturbing Cosine Router in Mixture of Experts

Huy Nguyen, Pedram Akbarian, Trang Pham +3

The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable…

stat.ML2024

Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?

Huy Nguyen, Pedram Akbarian, Nhat Ho

Dense-to-sparse gating mixture of experts (MoE) has recently become an effective alternative to a well-known sparse MoE. Rather than fixing the number of activated experts as in th…