12 citations · 17 across the 4 of their papers we have counts for
7 papers · 1 filter
MAP Estimation with Denoisers: Convergence Rates and Guarantees
Scott Pesme, Giacomo Meanti, Michael Arbel +1
Denoiser models have become powerful tools for inverse problems, enabling the use of pretrained networks to approximate the score of a smoothed prior distribution. These models are…
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir
We study the dynamics of gradient flow with small weight decay on general training losses . Under mild regularity assumptions and assuming convergen…
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
Hristo Papazov, Scott Pesme, Nicolas Flammarion
In this work, we investigate the effect of momentum on the optimisation trajectory of gradient descent. We leverage a continuous-time approach in the analysis of momentum gradient…
Saddle-to-Saddle Dynamics in Diagonal Linear Networks
Scott Pesme, Nicolas Flammarion
In this paper we fully describe the trajectory of gradient flow over diagonal linear networks in the limit of vanishing initialisation. We show that the limiting flow successively…
(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability
Mathieu Even, Scott Pesme, Suriya Gunasekar +1
In this paper, we investigate the impact of stochasticity and large stepsizes on the implicit regularisation of gradient descent (GD) and stochastic gradient descent (SGD) over dia…
On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent
Scott Pesme, Aymeric Dieuleveut, Nicolas Flammarion
Constant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase d…