Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
LOTION: Smoothing the Optimization Landscape for Quantized Training
Mujin Kwun, Depen Morwani, Chloe Huangyuan Su +3
Optimizing neural networks for quantized objectives is fundamentally challenging because the quantizer is piece-wise constant, yielding zero gradients everywhere except at quantiza…
cs.LG2025
Characterization and Mitigation of Training Instabilities in Microscaling Formats
Huangyuan Su, Mujin Kwun, Stephanie Gil +2
Training large language models is an expensive, compute-bound process that must be repeated as models scale, algorithms improve, and new data is collected. To address this, next-ge…