1 paper
Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho +4
Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities pro…