3 papers
cs.CV2026
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
Sayeh Sharify, Mahsa Salmani, Hesham Mostafa
Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While aggressive Post-Training Qu…
cs.CL2025
LLM Inference Acceleration via Efficient Operation Fusion
Mahsa Salmani, Ilya Soloveychik
The rapid development of the Transformer-based Large Language Models (LLMs) in recent years has been closely linked to their ever-growing and already enormous sizes. Many LLMs cont…
cs.LG2024
SLaNC: Static LayerNorm Calibration
Mahsa Salmani, Nikita Trukhanov, Ilya Soloveychik
The ever increasing sizes of Large Language Models (LLMs) beyond hundreds of billions of parameters have generated enormous pressure on the manufacturers of dedicated hardware acce…