6 papers
Deriving Neural Scaling Laws from the statistics of natural language
Francesco Cagnetta, Allan Raventós, Surya Ganguli +1
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict t…
Deep networks learn to parse uniform-depth context-free languages from local statistics
Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal repres…
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
Alessandro Favero, Antonio Sclocchi, Francesco Cagnetta +2
Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a co…
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
Francesco Cagnetta, Alessandro Favero, Antonio Sclocchi +1
How do neural language models acquire a language's structure when trained for next-token prediction? We address this question by deriving theoretical scaling laws for neural networ…
Learning curves theory for hierarchically compositional data with power-law distributed features
Francesco Cagnetta, Hyunmo Kang, Matthieu Wyart
Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data…
Towards a theory of how the structure of language is acquired by deep neural networks
Francesco Cagnetta, Matthieu Wyart
How much data is required to learn the structure of a language via next-token prediction? We study this question for synthetic datasets generated via a Probabilistic Context-Free G…