3 papers
cs.LG2025
Early Attentive Sparsification Accelerates Neural Speech Transcription
Zifei Xu, Sayeh Sharify, Hesham Mostafa +3
Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…
cs.LG2024
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
Zifei Xu, Sayeh Sharify, Wanzin Yazar +2
Large language models of high parameter counts are computationally expensive, yet can be made much more efficient by compressing their weights to very low numerical precision. This…
cs.LG2024
Scaling Laws for Post Training Quantized Large Language Models
Zifei Xu, Alexander Lan, Wanzin Yazar +3
Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling…