additive-multiplicative updates 1large language models 1low-precision training 1optimizer design 1quantized neural networks 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
Xiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup +1
The paper introduces M+Adam, an optimizer that blends additive and multiplicative updates to enable stable low‑precision training of large language models without keeping high‑prec…
cs.LG2025
TensorGRaD: Tensor Gradient Robust Decomposition for Memory-Efficient Neural Operator Training
Sebastian Loeschcke, David Pitt, Robert Joseph George +5
Scientific problems require resolving multi-scale phenomena across different resolutions and learning solution operators in infinite-dimensional function spaces. Neural operators p…