2 papers
cs.LG2025
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
Alessandro Favero, Antonio Sclocchi, Matthieu Wyart
Diffusion probabilistic models have become a cornerstone of modern generative AI, yet the mechanisms underlying their generalization remain poorly understood. In fact, if these mod…
cs.LG2025
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
Francesco Cagnetta, Alessandro Favero, Antonio Sclocchi +1
How do neural language models acquire a language's structure when trained for next-token prediction? We address this question by deriving theoretical scaling laws for neural networ…