9 papers
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
Hai Hu, Siyuan Song, Chongtian Shao +3
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to a…
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Siyuan Song, Zhiheng Qian, Yunhao Zhang +11
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences
Sriram Padmanabhan, Siyuan Song, Kanishka Misra
Language places subtle constraints on how we make inductive inferences. Developmental evidence by Gelman et al. (2002) has shown children (4 years and older) to differentiate among…
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
Yikang Liu, Yeting Shen, Hongao Zhu +9
We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train…
What Can String Probability Tell Us About Grammaticality?
Jennifer Hu, Ethan Gotlieb Wilcox, Siyuan Song +2
What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammatic…