activity
20242026
most citedLanguage Models Fail to Introspect About Their Knowledge of Language

2 citations · 3 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL2026

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

Siyuan Song, Zhiheng Qian, Yunhao Zhang +11

This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…

cs.CL2026

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

Hai Hu, Siyuan Song, Chongtian Shao +3

In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to a…

cs.CL2026

Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences

Sriram Padmanabhan, Siyuan Song, Kanishka Misra

Language places subtle constraints on how we make inductive inferences. Developmental evidence by Gelman et al. (2002) has shown children (4 years and older) to differentiate among…

cs.CL2025

What Can String Probability Tell Us About Grammaticality?

Jennifer Hu, Ethan Gotlieb Wilcox, Siyuan Song +2

What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammatic…

cs.CL2025

BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data

Jaap Jumelet, Abdellah Fourtassi, Akari Haga +23

We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally pla…

cs.CL2025

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…