5 papers
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Siyuan Song, Zhiheng Qian, Yunhao Zhang +11
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…
LLMs for automatic annotation of Mandarin narrative transcripts
Qingwen Zhao, Hongao Zhu, Yunqi He +3
Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consum…
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
Yikang Liu, Yeting Shen, Hongao Zhu +9
We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train…
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
Byung-Doh Oh, Hongao Zhu, William Schuler
In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been sp…
Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled
Yi-Chien Lin, Hongao Zhu, William Schuler
The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive 'quality-power'…