1 paper
Haochen Huang, Shuzhang Zhong, Zhe Zhang +5
Large Language Models (LLMs) with Mixture-of-Expert (MoE) architectures achieve superior model performance with reduced computation costs, but at the cost of high memory capacity a…