7 papers
How isotropic kernels perform on simple invariants
Jonas Paccolat, Stefano Spigler, Matthieu Wyart
We investigate how the training curve of isotropic kernel methods depends on the symmetry of the task to be learned, in several settings. (i) We consider a regression task, where t…
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot +1
Two distinct limits for deep learning have been derived as the network width , depending on how the weights of the last layer scale with . In the Neural Tan…
Asymptotic learning curves of kernel methods: empirical data v.s. Teacher-Student paradigm
Stefano Spigler, Mario Geiger, Matthieu Wyart
How many training data are needed to learn a supervised task? It is often observed that the generalization error decreases as where is the number of training examples…
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler +6
Supervised deep learning involves the training of neural networks with a large number of parameters. For large enough , in the so-called over-parametrized regime, one can es…
A jamming transition from under- to over-parametrization affects loss landscape and generalization
Stefano Spigler, Mario Geiger, Stéphane d'Ascoli +3
We argue that in fully-connected networks a phase transition delimits the over- and under-parametrized regimes where fitting can or cannot be achieved. Under some general condition…
The jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d'Ascoli +4
Deep learning has been immensely successful at a variety of tasks, ranging from classification to AI. Learning corresponds to fitting training data, which is implemented by descend…