1 paper
Yujie Zhang, Shivam Aggarwal, Tulika Mitra
Mixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory-constrained devices. While GPUs offer…