3 papers
cs.LG2026
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
Maximilian Kleinegger, Elvir Crnčević, Dan Alistarh
Matryoshka Quantization (MatQuant) is a recent quantization approach showing that a single integer-quantized model can be served across multiple precisions, by slicing the most sig…
cs.LG2026
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
Eugenia Iofinova, Dan Alistarh
As artificial neural networks, and specifically large language models, have improved rapidly in capabilities and quality, they have increasingly been deployed in real-world applica…
cs.CL2026
ECO: Quantized Training without Full-Precision Master Weights
Mahdi Nikdan, Amir Zandieh, Dan Alistarh +1
Quantization has significantly improved the compute and memory efficiency of Large Language Model (LLM) training. However, existing approaches still rely on accumulating their upda…