4 papers
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
Alireza Dadgarnia, Soroush Tabesh, Mahdi Nikdan +4
Quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-3 bits per parameter. The state of…
Model Compression with Exact Budget Constraints via Riemannian Manifolds
Michael Helcig, Dan Alistarh
Assigning one of K options to each of N groups under a total cost budget is a recurring problem in efficient AI, including mixed-precision quantization, non-uniform pruning, and ex…
Statistically-Lossless Quantization of Large Language Models
Michael Helcig, Eldar Kurtic, Dan Alistarh
Model quantization has become essential for efficient large language model deployment, yet existing approaches present clear trade-offs: methods such as GPTQ and AWQ achieve practi…
FedCCL: Federated Clustered Continual Learning Framework for Privacy-focused Energy Forecasting
Michael A. Helcig, Stefan Nastic
Privacy-preserving distributed model training is crucial for modern machine learning applications, yet existing Federated Learning approaches struggle with heterogeneous data distr…