collaborators

7 papers

cs.LG2026

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Yann Bouquet, Alireza Khodamoradi, Kristof Denolf +1

Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-b…

cs.DC2026

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

Victor J. B. Jung, Gagandeep Singh, Joseph Melber +3

The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-ch…

cs.LG2026

AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation

Seonggon Kim, Alireza Khodamoradi, Pranathi Vasireddy +2

Hadamard transforms have become a key tool for stabilizing low-precision training, but existing methods apply them uniformly across tensors and computation paths. We show that this…

cs.LG2026

LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization

Yann Bouquet, Alireza Khodamoradi, Sophie Yáng Shen +2

Post-training quantization (PTQ) is essential for deploying large diffusion transformers on resource-constrained hardware, but aggressive 4-bit quantization significantly degrades…

cs.AR2025

Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs

Endri Taka, Andre Roesti, Joseph Melber +3

The high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen…

cs.CL2025

From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR

Erwei Wang, Samuel Bayliss, Andra Bisca +19

General-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures…