activity
20212026
most citedFindings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

72 citations · 85 across the 19 of their papers we have counts for

collaborators
Showing cs.CLShow all

31 papers · 1 filter

cs.CL2026

How Anthropomorphic Language Impacts Public Perceptions of AI

Betty Li Hou, Sophie Hao, Sunoo Park +1

Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and characteristics to the system. This practic…

cs.CL2026

Simulating Human Memory with Language Models

Qihan Wang, Nicholas Tomlin, Michael Hu +2

Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic m…

cs.CL2026

Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

William Timkey, Brian Dillon, Tal Linzen

Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and nex…

cs.CL2026

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

Michael Y. Hu, Apurva Gandhi, Kyunghyun Cho +2

Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretraining, data composition is a key d…

cs.CL2026

Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction

Jackson Petty, Jaulie Goe, Tal Linzen

Low-resource languages pose a challenge for machine translation with large language models (LLMs), which require large amounts of training data. One potential way to circumvent thi…

cs.CL2026

BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop

Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul +7

The goal of the BabyLM is to stimulate new research connections between cognitive modeling and language model pretraining. We invite contributions in this vein to the BabyLM Worksh…