collaborators

7 papers

cs.LG2026

A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts

Viet Nguyen, Tuan Minh Pham, Thinh Cao +4

Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to impro…

stat.ML2026

Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function

Tuan Minh Pham, Thinh Cao, Viet Nguyen +3

The sigmoid gate in mixture-of-experts (MoE) models has been empirically shown to outperform the softmax gate across several tasks, ranging from approximating feed-forward networks…

stat.ME2025

Conformalized Bayesian Inference, with Applications to Random Partition Models

Nicola Bariletto, Nhat Ho, Alessandro Rinaldo

Bayesian posterior distributions naturally represent parameter uncertainty informed by data. However, when the parameter space is complex, as in many nonparametric settings where i…

stat.ML2025

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

Fanqi Yan, Huy Nguyen, Dung Le +3

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstr…

cs.LG2025

On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating

Huy Nguyen, Thong T. Doan, Quang Pham +3

Mixture of experts (MoE) methods are a key component in most large language model architectures, including the recent series of DeepSeek models. Compared to other MoE implementatio…

stat.ML2025

Convergence Rates for Softmax Gating Mixture of Experts

Huy Nguyen, Nhat Ho, Alessandro Rinaldo

Mixture of experts (MoE) has recently emerged as an effective framework to advance the efficiency and scalability of machine learning models by softly dividing complex tasks among…