collaborators

6 papers

cs.IR2026

Token-Level Credit Assignment Optimization for Generative Document Retrieval

Xinpeng Zhao, Yang Liu, Ran Chen +6

Generative retrieval models perform document retrieval by autoregressively generating document identifiers (DocIDs). This process naturally forms a sequential decision problem, i.e…

cs.CL2026

Rethinking Data Mixing from the Perspective of Large Language Models

Yuanjian Xu, Tianze Sun, Changwei Xu +7

Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly reduce generalization. Although…

cs.IR2026

DiffuGR: Generative Document Retrieval with Diffusion Language Models

Xinpeng Zhao, Zhaochun Ren, Yukun Zhao +9

Generative retrieval (GR) reframes document retrieval as an end-to-end task of generating sequential document identifiers (DocIDs). Existing GR methods predominantly rely on left-t…

cs.CL2025

Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps

Yijiong Yu, Yongfeng Huang, Zhixiao Qi +4

Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long…

cs.SE2025

CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation

Wenjing Yin, Tianze Sun, Yijiong Yu +19

Large language models (LLMs) play a crucial role in software engineering, excelling in tasks like code generation and maintenance. However, existing benchmarks are often narrow in…

cs.CL2025

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Yijiong Yu, Ziyun Dai, Zekun Wang +3

Large language models (LLMs) have demonstrated remarkable capabilities, but their success heavily relies on the quality of pretraining corpora. For Chinese LLMs, the scarcity of hi…