4 papers
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Siyuan Song, Zhiheng Qian, Yunhao Zhang +11
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
Zhiheng Qian, Aini Li, Hai Hu +1
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by tr…
Math Natural Language Inference: this should be easy!
Valeria de Paiva, Qiyue Gao, Hai Hu +4
We ask whether contemporary LLMs are able to perform natural language inference (NLI) tasks on mathematical texts. We call this the Math NLI problem. We construct a corpus of Math…
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
Yikang Liu, Yeting Shen, Hongao Zhu +9
We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train…