Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
Sohir Maskey, Constantin Eichenberg, Johannes Messner +1
Quantization-aware training (QAT) is an effective method to drastically reduce the memory footprint of LLMs while keeping performance degradation at an acceptable level. However, t…
cs.LG2025
u-P: The Unit-Scaled Maximal Update Parametrization
Charlie Blake, Constantin Eichenberg, Josef Dean +7
The Maximal Update Parametrization (P) aims to make the optimal hyperparameters (HPs) of a model independent of its size, allowing them to be swept using a cheap proxy model ra…