5 papers
Early Attentive Sparsification Accelerates Neural Speech Transcription
Zifei Xu, Sayeh Sharify, Hesham Mostafa +3
Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
Zifei Xu, Sayeh Sharify, Wanzin Yazar +2
Large language models of high parameter counts are computationally expensive, yet can be made much more efficient by compressing their weights to very low numerical precision. This…
Scaling Laws for Post Training Quantized Large Language Models
Zifei Xu, Alexander Lan, Wanzin Yazar +3
Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling…
Post Training Quantization of Large Language Models with Microscaling Formats
Sayeh Sharify, Utkarsh Saxena, Zifei Xu +3
Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage…
Self-Selected Attention Span for Accelerating Large Language Model Inference
Tian Jin, Wanzin Yazar, Zifei Xu +2
Large language models (LLMs) can solve challenging tasks. However, their inference computation on modern GPUs is highly inefficient due to the increasing number of tokens they must…