1 citations · 1 across the 2 of their papers we have counts for
3 papers
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Mike Lasby, Ivan Lazarevich, Nish Sinnadurai +3
Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating res…
QGen: On the Ability to Generalize in Quantization Aware Training
MohammadHossein AskariHemmat, Ahmadreza Jeddi, Reyhane Askari Hemmat +6
Quantization lowers memory usage, computational requirements, and latency by utilizing fewer bits to represent model weights and activations. In this work, we investigate the gener…
Accelerating Deep Neural Networks via Semi-Structured Activation Sparsity
Matteo Grimaldi, Darshan C. Ganji, Ivan Lazarevich +1
The demand for efficient processing of deep neural networks (DNNs) on embedded devices is a significant challenge limiting their deployment. Exploiting sparsity in the network's fe…