1 paper · 1 filter
Zeyu Zhu, Gang Li, Peisong Wang +5
Mixture of Experts (MoE) architectures significantly enhance the capacity of LLMs without proportional increases in computation, but at the cost of a vast parameter size. Offloadin…