additive-multiplicative updates 1large language models 1low-precision training 1optimizer design 1quantized neural networks 1
From the 1 of 2 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
Xiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup +1
The paper introduces M+Adam, an optimizer that blends additive and multiplicative updates to enable stable low‑precision training of large language models without keeping high‑prec…
cs.LG2024
LoQT: Low-Rank Adapters for Quantized Pretraining
Sebastian Loeschcke, Mads Toftrup, Michael J. Kastoryano +2
Despite advances using low-rank adapters and quantization, pretraining of large models on consumer hardware has not been possible without model sharding, offloading during training…