5 papers
Incremental Learning in Mirror Flows
Raphaël Berthier, Loucas Pillaud-Vivien
We study mirror flows generated by a convex quadratic loss and a general convex lower semicontinuous mirror potential. We show that, when initialized near the boundary of the domai…
Diagonal Linear Networks and the Lasso Regularization Path
Raphaël Berthier
Diagonal linear networks are neural networks with linear activation and diagonal weight matrices. Their theoretical interest is that their implicit regularization can be rigorously…
Acceleration of Gossip Algorithms through the Euler-Poisson-Darboux Equation
Raphaël Berthier, Mufan Bill Li
Gossip algorithms and their accelerated versions have been studied exclusively in discrete time on graphs. In this work, we take a different approach, and consider the scaling limi…
Learning time-scales in two-layers neural networks
Raphaël Berthier, Andrea Montanari, Kangjie Zhou
Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averagi…
Attention layers provably solve single-location regression
Pierre Marion, Raphaël Berthier, Gérard Biau +1
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal li…