3 papers
cs.LG2026
Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization
Etienne Boursier, Matthew Bowditch, Matthias Englert +1
The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint. While weight decay is standard practice in modern training procedure…
cs.LG2026
Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
James Town, Etienne Boursier, Ben Lewis +2
The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is…
cs.LG2025
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir
We study the dynamics of gradient flow with small weight decay on general training losses . Under mild regularity assumptions and assuming convergen…