6 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CL2024
MH-MoE: Multi-Head Mixture-of-Experts
Shaohan Huang, Xun Wu, Shuming Ma +1
Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces…
cs.CL2024★ 6 cited
Multi-Head Mixture-of-Experts
Xun Wu, Shaohan Huang, Wenhui Wang +1
Sparse Mixtures of Experts (SMoE) scales model capacity without significant increases in training and inference costs, but exhibits the following two issues: (1) Low expert activat…
cs.CL2024★ 3 cited
Mixture of LoRA Experts
Xun Wu, Shaohan Huang, Furu Wei
LoRA has gained widespread acceptance in the fine-tuning of large pre-trained models to cater to a diverse array of downstream tasks, showcasing notable effectiveness and efficienc…