4 papers
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23
We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…
Propositional Logic for Probing Generalization in Neural Networks
Anna Langedijk, Jaap Jumelet, Willem Zuidema
The extent to which neural networks are able to acquire and represent symbolic rules remains a key topic of research and debate. Much current work focuses on the impressive capabil…
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
Ezgi Başar, Francesca Padovani, Jaap Jumelet +1
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs).…
Finding Structure in Language Models
Jaap Jumelet
When we speak, write or listen, we continuously make predictions based on our knowledge of a language's grammar. Remarkably, children acquire this grammatical knowledge within just…