2 papers
cs.LG2026
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL
Juliette Decugis, Sean O'Brien, Francis Bach +2
Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse. Advantage functions offer an appealing fix:…
cs.LG2026
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
Sacchit Kale, Piyushi Manupriya, Pierre Marion +2
Gradient descent and stochastic gradient descent are central to modern machine learning, yet their behavior under large step sizes remains theoretically unclear. Recent work sugges…