2 papers
cs.LG2025
KurTail : Kurtosis-based LLM Quantization
Mohammad Sadegh Akhondzadeh, Aleksandar Bojchevski, Evangelos Eleftheriou +1
One of the challenges of quantizing a large language model (LLM) is the presence of outliers. Outliers often make uniform quantization schemes less effective, particularly in extre…
cs.LG2024
EfQAT: An Efficient Framework for Quantization-Aware Training
Saleh Ashkboos, Bram Verhoef, Torsten Hoefler +2
Quantization-aware training (QAT) schemes have been shown to achieve near-full precision accuracy. They accomplish this by training a quantized model for multiple epochs. This is c…