2 papers
cs.LG2026
Loss Smoothing for Stable Adaptation Under Distribution Shift
Darshan Patil, Ekaterina Lobacheva, Razvan Pascanu +1
In settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation methods typically optimize the target ob…
cs.LG2025
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan +4
This work aims to understand how scaling improves language models, specifically in terms of training dynamics. We find that language models undergo loss deceleration early in train…