23 papers
Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts
Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen +4
Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networ…
On the Geometry of Separation in Finite Gaussian Mixtures
Huy Nguyen, Dung Le, Alessandro Rinaldo +1
We study an open problem of understanding the effects of the minimum component separation on the convergence rates of parameter estimation in finite Gaussian mixtures. We address t…
On Bayesian Softmax-Gated Mixture-of-Experts Models
Nicola Bariletto, Huy Nguyen, Nhat Ho +1
Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert models through an input-dependent…
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
Minh Le, Anh Nguyen, Huy Nguyen +3
Visual Prompt Tuning (VPT) has proven effective for parameter-efficient adaptation of pre-trained vision models to downstream tasks by inserting task-specific learnable prompt toke…
DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
Nghiem T. Diep, Hien Dang, Tuan Truong +3
Parameter-efficient fine-tuning (PEFT) methods have become the standard paradigm for adapting large-scale models. Among these techniques, Weight-Decomposed Low-Rank Adaptation (DoR…
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
Viet Nguyen, Tuan Minh Pham, Thinh Cao +4
Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to impro…