most citedFindings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

3 citations · 3 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CL2025

BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data

Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23

We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…

cs.CL2025

The Harmonic Structure of Information Contours

Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak +5

The uniform information density (UID) hypothesis proposes that speakers aim to distribute information evenly throughout a text, balancing production effort and listener comprehensi…

cs.CL2025

Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models

Lennart Stöpler, Rufat Asadli, Mitja Nikolaus +2

We propose a method for training language models in an interactive setting inspired by child language acquisition. In our setting, a speaker attempts to communicate some informatio…

cs.CL2025

BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop

Lucas Charpentier, Leshem Choshen, Ryan Cotterell +11

BabyLM aims to dissolve the boundaries between cognitive modeling and language modeling. We call for both workshop papers and for researchers to join the 3rd BabyLM competition. As…

cs.CL2025

Can Language Models Learn Typologically Implausible Languages?

Tianyang Xu, Tatsuki Kuribayashi, Yohei Oseki +2

Grammatical features across human languages show intriguing correlations often attributed to learning biases in humans. However, empirical evidence has been limited to experiments…

cs.CL2025

A Distributional Perspective on Word Learning in Neural Language Models

Filippo Ficarra, Ryan Cotterell, Alex Warstadt

Language models (LMs) are increasingly being studied as models of human language learners. Due to the nascency of the field, it is not well-established whether LMs exhibit similar…