72 citations · 85 across the 19 of their papers we have counts for
31 papers · 1 filter
How Anthropomorphic Language Impacts Public Perceptions of AI
Betty Li Hou, Sophie Hao, Sunoo Park +1
Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and characteristics to the system. This practic…
Simulating Human Memory with Language Models
Qihan Wang, Nicholas Tomlin, Michael Hu +2
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic m…
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
William Timkey, Brian Dillon, Tal Linzen
Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and nex…
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
Michael Y. Hu, Apurva Gandhi, Kyunghyun Cho +2
Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretraining, data composition is a key d…
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
Jackson Petty, Jaulie Goe, Tal Linzen
Low-resource languages pose a challenge for machine translation with large language models (LLMs), which require large amounts of training data. One potential way to circumvent thi…
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul +7
The goal of the BabyLM is to stimulate new research connections between cognitive modeling and language model pretraining. We invite contributions in this vein to the BabyLM Worksh…