9 papers · 1 filter
CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
Francesca Padovani, Xiulin Yang, Bastian Bunzeck +4
CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of t…
Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition
Francesca Padovani, Jaap Jumelet, Yevgen Matusevych +1
Is child-directed language (CDL) optimized to support language learning, and which aspects of linguistic development does it facilitate? We investigate this question using neural l…
Vocabulary shapes cross-lingual variation of word-order learnability in language models
Jonas Mayer Martins, Jaap Jumelet, Viola Priesemann +1
Why do some languages like Czech permit free word order, while others like English do not? We address this question by pretraining transformer language models on a spectrum of synt…
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul +7
The goal of the BabyLM is to stimulate new research connections between cognitive modeling and language model pretraining. We invite contributions in this vein to the BabyLM Worksh…
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23
We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
Ezgi Başar, Francesca Padovani, Jaap Jumelet +1
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs).…