1 citations · 1 across the 2 of their papers we have counts for
17 papers
Deriving Neural Scaling Laws from the statistics of natural language
Francesco Cagnetta, Allan Raventós, Surya Ganguli +1
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict t…
Symmetry in language statistics shapes the geometry of model representations
Dhruva Karkada, Daniel J. Korchinski, Andres Nava +2
The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth on…
Deep networks learn to parse uniform-depth context-free languages from local statistics
Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal repres…
Learn from your own latents and not from tokens: A sample-complexity theory
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological lea…
Sampling Data with Chains of Forward-Backward Diffusion Steps
Hyunmo Kang, Noam Itzhak Levi, Corinna Elena Wegner +2
Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward…
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
Andres Nava, Matthieu Wyart
We propose a distributional theory of how hypernymy -- the ``is-a'' relation between general and specific concepts -- is encoded geometrically in language representations. Starting…