2 papers
cs.AI2025
QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads
Muhammad Ahmad, Khurram Mazher, Saqib Akram +2
We present QuantX: a tailored suite of recipes for LLM and VLM quantization. It is capable of quantizing down to 3-bit resolutions with minimal loss in performance. The quantizatio…
cs.AR2025
Accelerating GenAI Workloads by Enabling RISC-V Microkernel Support in IREE
Adeel Ahmad, Ahmad Tameem Kamal, Nouman Amir +2
This project enables RISC-V microkernel support in IREE, an MLIR-based machine learning compiler and runtime. The approach begins by enabling the lowering of MLIR linalg dialect co…