7 papers
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
Siyuan Song, Zhiheng Qian, Yunhao Zhang +11
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 10…
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
Chunyu Ye, Yunhao Zhang, Jingyuan Sun +3
Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecti…
Computational Lesions in Multilingual Language Models Separate Shared and Language-specific Brain Alignment
Yang Cui, Jingyuan Sun, Yizheng Sun +8
How the brain supports language across different languages is a basic question in neuroscience and a useful test for multilingual artificial intelligence. Neuroimaging has identifi…
Component-Level Lesioning of Language Models Reveals Clinically Aligned Aphasia Phenotypes
Yifan Wang, Jichen Zheng, Jingyuan Sun +5
Large language models (LLMs) increasingly exhibit human-like linguistic behaviors and internal representations that they could serve as computational simulators of language cogniti…
Discovering Semantic Subdimensions through Disentangled Conceptual Representations
Yunhao Zhang, Shaonan Wang, Nan Lin +3
Understanding the core dimensions of conceptual semantics is fundamental to uncovering how meaning is organized in language and the brain. Existing approaches often rely on predefi…
Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
Yifan Wang, Jingyuan Sun, Jichen Zheng +5
The striking alignment between large language models (LLMs) and human brain activity positions them as powerful models of healthy cognition. This parallel raises a fundamental ques…