6 papers
Tight Clusters Make Specialized Experts
Stefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev +1
Sparse Mixture-of-Experts (MoE) architectures have emerged as a promising approach to decoupling model capacity from computational cost. At the core of the MoE model is the router,…
Almost Asymptotically Optimal Active Clustering Through Pairwise Observations
Rachel S. Y. Teo, P. N. Karthik, Ramya Korlakai Vinayak +1
We propose a new analysis framework for clustering items into an unknown number of distinct groups using noisy and actively collected responses. At each time step, an agent…
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
Minh-Khoi Nguyen-Nhat, Rachel S. Y. Teo, Laziz Abdullaev +3
Sparse Mixture of Experts (SMoE) has emerged as a promising solution to achieving unparalleled scalability in deep learning by decoupling model parameter count from computational c…
The Blessing and Curse of Dimensionality in Safety Alignment
Rachel S. Y. Teo, Laziz U. Abdullaev, Tan M. Nguyen
The focus on safety alignment in large language models (LLMs) has increased significantly due to their widespread adoption across different domains. The scale of LLMs play a contri…
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
Rachel S. Y. Teo, Tan M. Nguyen
Large-scale pre-training of deep models, followed by fine-tuning them, has become the cornerstone of natural language processing (NLP). The prevalence of data coupled with computat…
CAMEx: Curvature-aware Merging of Experts
Dung V. Nguyen, Minh H. Nguyen, Luc Q. Nguyen +3
Existing methods for merging experts during model training and fine-tuning predominantly rely on Euclidean geometry, which assumes a flat parameter space. This assumption can limit…