2 papers
stat.ML2025
Feature learning from non-Gaussian inputs: the case of Independent Component Analysis in high dimensions
Fabiola Ricci, Lorenzo Bardone, Sebastian Goldt
Deep neural networks learn structured features from complex, non-Gaussian inputs, but the mechanisms behind this process remain poorly understood. Our work is motivated by the obse…
cs.CL2024
A distributional simplicity bias in the learning dynamics of transformers
Riccardo Rende, Federica Gerace, Alessandro Laio +1
The remarkable capability of over-parameterised neural networks to generalise effectively has been explained by invoking a ``simplicity bias'': neural networks prevent overfitting…