collaborators

5 papers

cs.LG2025

Early Attentive Sparsification Accelerates Neural Speech Transcription

Zifei Xu, Sayeh Sharify, Hesham Mostafa +3

Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…

cs.LG2025

Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs

Zifei Xu, Sayeh Sharify, Wanzin Yazar +2

Large language models of high parameter counts are computationally expensive, yet can be made much more efficient by compressing their weights to very low numerical precision. This…

cs.LG2024

Scaling Laws for Post Training Quantized Large Language Models

Zifei Xu, Alexander Lan, Wanzin Yazar +3

Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling…

cs.LG2024

Post Training Quantization of Large Language Models with Microscaling Formats

Sayeh Sharify, Utkarsh Saxena, Zifei Xu +3

Large Language Models (LLMs) have distinguished themselves with outstanding performance in complex language modeling tasks, yet they come with significant computational and storage…

cs.CL2024

Self-Selected Attention Span for Accelerating Large Language Model Inference

Tian Jin, Wanzin Yazar, Zifei Xu +2

Large language models (LLMs) can solve challenging tasks. However, their inference computation on modern GPUs is highly inefficient due to the increasing number of tokens they must…