activity
20212026
most citedFindings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

72 citations · 78 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

38 papers · 1 filter

cs.CL2026

A Formal Limitation on Learning Human Language From Textual Corpora

Emily Cheng, Ryan Cotterell

Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of te…

cs.CL2026

Surprisal Theory is Tautological (without Rational Grounding)

Ryan Cotterell

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is…

cs.CL202572 cited

Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Alex Warstadt, Aaron Mueller, Leshem Choshen +8

Children can acquire language from less than 100 million words of input. Large language models are far less data-efficient: they typically require 3 or 4 orders of magnitude more d…

cs.CL20243 cited

Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Michael Y. Hu, Aaron Mueller, Candace Ross +7

The BabyLM Challenge is a community effort to close the data-efficiency gap between human and computational language learners. Participants compete to optimize language model train…

cs.CL2024

Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse

Eleftheria Tsipidi, Franz Nowak, Ryan Cotterell +3

The Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication. Of course, i…

cs.CL2024

Reverse-Engineering the Reader

Samuel Kiegeland, Ethan Gotlieb Wilcox, Afra Amini +2

Numerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition. In this paper…