Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
Sohir Maskey, Constantin Eichenberg, Johannes Messner +1
Quantization-aware training (QAT) is an effective method to drastically reduce the memory footprint of LLMs while keeping performance degradation at an acceptable level. However, t…
cs.LG2024
Approximate Top- for Increased Parallelism
Oscar Key, Luka Ribar, Alberto Cattaneo +2
We present an evaluation of bucketed approximate top- algorithms. Computing top- exactly suffers from limited parallelism, because the largest values must be aggregated a…