collaborators

12 papers

cs.CL2026

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

Qingjie Zhang, Xingzhang Ren, Zixuan Chen +6

Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific…

cs.CL2026

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

Shaobo Wang, Guo Chen, Ziyue Wang +5

With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The challenge is that base pre-…

cs.LG2026

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL

Shaobo Wang, Yujie Chen, Yafeng Sun +9

Modern reasoning agents are increasingly evaluated on their ability to generate multiple valid solution paths, plans, or tool-use traces for a given input. Standard reward-maximizi…

cs.CL2026

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

Shaobo Wang, Xuan Ouyang, Tianyi Xu +9

As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall, pre-training is shifting from more tokens to better tokens. However, existing methods either…

cs.CL2025

Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?

Shaobo Wang, Cong Wang, Wenjie Fu +11

As the demand for comprehensive evaluations of diverse model capabilities steadily increases, benchmark suites have correspondingly grown significantly in scale. Despite notable ad…

cs.CL2025

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…