1 citations · 1 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026★ 1 cited
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Mike Lasby, Ivan Lazarevich, Nish Sinnadurai +3
Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating res…
cs.LG2024
QGen: On the Ability to Generalize in Quantization Aware Training
MohammadHossein AskariHemmat, Ahmadreza Jeddi, Reyhane Askari Hemmat +6
Quantization lowers memory usage, computational requirements, and latency by utilizing fewer bits to represent model weights and activations. In this work, we investigate the gener…