collaborators

5 papers

cs.CL2026

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

Siyuan Song, Zhiheng Qian, Yunhao Zhang +11

This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…

cs.CL2026

LLMs for automatic annotation of Mandarin narrative transcripts

Qingwen Zhao, Hongao Zhu, Yunqi He +3

Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consum…

cs.CL2025

A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

Yikang Liu, Yeting Shen, Hongao Zhu +9

We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train…

cs.CL2025

The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage

Byung-Doh Oh, Hongao Zhu, William Schuler

In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been sp…

cs.CL2025

Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled

Yi-Chien Lin, Hongao Zhu, William Schuler

The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive 'quality-power'…