Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Scale Weight Decay and Train Better
Anuj Apte
The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which caus…
cs.LG2026
Anytime Training with Schedule-Free Spectral Optimization
Anuj Apte, Pranav Deshpande, Niraj Kumar +2
Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading to strong path dependence and costly re-tuning as data availability changes. Sch…