3 citations · 3 across the 5 of their papers we have counts for
7 papers
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23
We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…
The Harmonic Structure of Information Contours
Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak +5
The uniform information density (UID) hypothesis proposes that speakers aim to distribute information evenly throughout a text, balancing production effort and listener comprehensi…
Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models
Lennart Stöpler, Rufat Asadli, Mitja Nikolaus +2
We propose a method for training language models in an interactive setting inspired by child language acquisition. In our setting, a speaker attempts to communicate some informatio…
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
Lucas Charpentier, Leshem Choshen, Ryan Cotterell +11
BabyLM aims to dissolve the boundaries between cognitive modeling and language modeling. We call for both workshop papers and for researchers to join the 3rd BabyLM competition. As…
Can Language Models Learn Typologically Implausible Languages?
Tianyang Xu, Tatsuki Kuribayashi, Yohei Oseki +2
Grammatical features across human languages show intriguing correlations often attributed to learning biases in humans. However, empirical evidence has been limited to experiments…
A Distributional Perspective on Word Learning in Neural Language Models
Filippo Ficarra, Ryan Cotterell, Alex Warstadt
Language models (LMs) are increasingly being studied as models of human language learners. Due to the nascency of the field, it is not well-established whether LMs exhibit similar…