6 papers
Token-Level Credit Assignment Optimization for Generative Document Retrieval
Xinpeng Zhao, Yang Liu, Ran Chen +6
Generative retrieval models perform document retrieval by autoregressively generating document identifiers (DocIDs). This process naturally forms a sequential decision problem, i.e…
Rethinking Data Mixing from the Perspective of Large Language Models
Yuanjian Xu, Tianze Sun, Changwei Xu +7
Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly reduce generalization. Although…
DiffuGR: Generative Document Retrieval with Diffusion Language Models
Xinpeng Zhao, Zhaochun Ren, Yukun Zhao +9
Generative retrieval (GR) reframes document retrieval as an end-to-end task of generating sequential document identifiers (DocIDs). Existing GR methods predominantly rely on left-t…
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
Yijiong Yu, Yongfeng Huang, Zhixiao Qi +4
Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long…
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
Wenjing Yin, Tianze Sun, Yijiong Yu +19
Large language models (LLMs) play a crucial role in software engineering, excelling in tasks like code generation and maintenance. However, existing benchmarks are often narrow in…
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
Yijiong Yu, Ziyun Dai, Zekun Wang +3
Large language models (LLMs) have demonstrated remarkable capabilities, but their success heavily relies on the quality of pretraining corpora. For Chinese LLMs, the scarcity of hi…