3 papers
cs.LG2026
Tight Clusters Make Specialized Experts
Stefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev +1
Sparse Mixture-of-Experts (MoE) architectures have emerged as a promising approach to decoupling model capacity from computational cost. At the core of the MoE model is the router,…
cs.LG2024
An Attention-based Framework for Fair Contrastive Learning
Stefan K. Nielsen, Tan M. Nguyen
Contrastive learning has proven instrumental in learning unbiased representations of data, especially in complex environments characterized by high-cardinality and high-dimensional…
cs.LG2024
Elliptical Attention
Stefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo +1
Pairwise dot-product self-attention is key to the success of transformers that achieve state-of-the-art performance across a variety of applications in language and vision. This do…