7 papers
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
Viet Nguyen, Tuan Minh Pham, Thinh Cao +4
Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to impro…
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
Tuan Minh Pham, Thinh Cao, Viet Nguyen +3
The sigmoid gate in mixture-of-experts (MoE) models has been empirically shown to outperform the softmax gate across several tasks, ranging from approximating feed-forward networks…
Conformalized Bayesian Inference, with Applications to Random Partition Models
Nicola Bariletto, Nhat Ho, Alessandro Rinaldo
Bayesian posterior distributions naturally represent parameter uncertainty informed by data. However, when the parameter space is complex, as in many nonparametric settings where i…
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
Fanqi Yan, Huy Nguyen, Dung Le +3
The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstr…
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating
Huy Nguyen, Thong T. Doan, Quang Pham +3
Mixture of experts (MoE) methods are a key component in most large language model architectures, including the recent series of DeepSeek models. Compared to other MoE implementatio…
Convergence Rates for Softmax Gating Mixture of Experts
Huy Nguyen, Nhat Ho, Alessandro Rinaldo
Mixture of experts (MoE) has recently emerged as an effective framework to advance the efficiency and scalability of machine learning models by softly dividing complex tasks among…