1 paper
Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun
Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processi…