4 papers
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
Viet Nguyen, Tuan Minh Pham, Thinh Cao +4
Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to impro…
Rethinking Multinomial Logistic Mixture of Experts with Sigmoid Gating Function
Tuan Minh Pham, Thinh Cao, Viet Nguyen +3
The sigmoid gate in mixture-of-experts (MoE) models has been empirically shown to outperform the softmax gate across several tasks, ranging from approximating feed-forward networks…
Reliably Detecting Model Failures in Deployment Without Labels
Viet Nguyen, Changjian Shui, Vijay Giri +4
The distribution of data changes over time; models operating in dynamic environments need retraining. But knowing when to retrain, without access to labels, is an open challenge si…
Diverse Prototypical Ensembles Improve Robustness to Subpopulation Shift
Minh Nguyen Nhat To, Paul F RWilson, Viet Nguyen +6
The subpopulationtion shift, characterized by a disparity in subpopulation distributibetween theween the training and target datasets, can significantly degrade the performance of…