4 papers
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Siyuan Song, Zhiheng Qian, Yunhao Zhang +11
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…
Creative Convergence or Imitation? Genre-Specific Homogeneity in LLM-Generated Chinese Literature
Yuanchi Ma, Kaize Shi, Hui He +5
Large Language Models (LLMs) have demonstrated remarkable capabilities in narrative generation. However, they often produce structurally homogenized stories, frequently following r…
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
Ziliang Qiu, Renfen Hu
The evaluation of LLMs' creativity represents a crucial research domain, though challenges such as data contamination and costly human assessments often impede progress. Drawing in…
Efficiently Building a Domain-Specific Large Language Model from Scratch: A Case Study of a Classical Chinese Large Language Model
Shen Li, Renfen Hu, Lijun Wang
General-purpose large language models demonstrate notable capabilities in language comprehension and generation, achieving results that are comparable to, or even surpass, human pe…