2 papers
cs.CV2026
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
Sayeh Sharify, Mahsa Salmani, Hesham Mostafa
Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While aggressive Post-Training Qu…
cs.LG2025
Early Attentive Sparsification Accelerates Neural Speech Transcription
Zifei Xu, Sayeh Sharify, Hesham Mostafa +3
Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…