1 paper · 1 filter
Francesco Cagnetta, Matthieu Wyart
How much data is required to learn the structure of a language via next-token prediction? We study this question for synthetic datasets generated via a Probabilistic Context-Free G…