1 paper · 1 filter
Nishant Gavhane, Arush Mehrotra, Rohit Chawla +1
The deployment of large-scale Mixture-of-Experts (MoE) models on edge devices presents significant challenges due to memory constraints. While MoE architectures enable efficient ut…