1 citations · 2 across the 5 of their papers we have counts for
4 papers · 1 filter
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Mike Lasby, Ivan Lazarevich, Nish Sinnadurai +3
Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating res…
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
Cem Üyük, Mike Lasby, Mohamed Yassin +2
Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices. Among existing compression approa…
Navigating Extremes: Dynamic Sparsity in Large Output Spaces
Nasib Ullah, Erik Schultheis, Mike Lasby +2
In recent years, Dynamic Sparse Training (DST) has emerged as an alternative to post-training pruning for generating efficient models. In principle, DST allows for a more memory ef…
Dynamic Sparse Training with Structured Sparsity
Mike Lasby, Anna Golubeva, Utku Evci +2
Dynamic Sparse Training (DST) methods achieve state-of-the-art results in sparse neural network training, matching the generalization of dense models while enabling sparse training…