2 papers
cs.LG2026
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
Ofir Gordon, Lior Dikstein, Arnon Netzer +2
Post-training quantization (PTQ) is a widely used approach for reducing the memory and compute costs of large language models (LLMs). Recent studies have shown that applying invert…
cs.LG2025
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
Ofir Gordon, Ariel Lapid, Elad Cohen +3
Deploying transformer-based neural networks on resource-constrained edge devices presents a significant challenge. This challenge is often addressed through various techniques, suc…