1 citations · 2 across the 8 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
math.OC2023
Leveraging the two timescale regime to demonstrate convergence of neural networks
Pierre Marion, Raphaël Berthier
We study the training dynamics of shallow neural networks, in a two-timescale regime in which the stepsizes for the inner layer are much smaller than those for the outer layer. In…
cs.LG2023
Learning time-scales in two-layers neural networks
Raphaël Berthier, Andrea Montanari, Kangjie Zhou
Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averagi…