activity
20242026
collaborators

5 papers

cs.CL2026

When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

Xu Guo, Jian Tong, Zhihui Lu +1

Synthetic data can be scaled along two routes: Source Expansion (SE), which enlarges the source by adding seed materials or generators, and Fixed-Source Synthesis (FSS), which hold…

cs.CL2026

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Xu Guo, Runyu Peng, Jian Tong +4

Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model perf…

cs.CL2026

Rethinking Multiple-Choice Questions for RLVR: Unlocking Potential via Distractor Design

Xu Guo, Qiming Ge, Jian Tong +8

Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capabilities of Large Language Models. When applied to RLVR, Multiple-Choice Questions (M…

cs.CL2025

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law

Qiming Ge, Shuhao Xing, Songyang Gao +8

Scaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different level…

cs.CL2024

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

Zhikai Lei, Tianyi Liang, Hanglei Hu +8

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However…