2 citations · 3 across the 4 of their papers we have counts for
1 paper · 1 filter
Bowen Pan, Yikang Shen, Haokun Liu +5
Mixture-of-Experts (MoE) language models can reduce computational costs by 2-4× compared to dense models without sacrificing performance, making them more efficient in compu…