activity
20242026
collaborators

9 papers

cs.CL2026

Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?

Hai Hu, Siyuan Song, Chongtian Shao +3

In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to a…

cs.CL2026

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Zhiheng Qian, Aini Li, Hai Hu +1

Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by tr…

cs.CL2026

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

Siyuan Song, Zhiheng Qian, Yunhao Zhang +11

This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…

cs.CL2026

LLMs for automatic annotation of Mandarin narrative transcripts

Qingwen Zhao, Hongao Zhu, Yunqi He +3

Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consum…

cs.CL2025

A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese

Yikang Liu, Yeting Shen, Hongao Zhu +9

We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the \textit{Ba} construction. We then train…

cs.CL2025

Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese

Yikang Liu, Wanyang Zhang, Yiming Wang +6

Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original…