Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Rethinking Data Mixing from the Perspective of Large Language Models
Yuanjian Xu, Tianze Sun, Changwei Xu +7
Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly reduce generalization. Although…
cs.CL2025
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
Yijiong Yu, Yongfeng Huang, Zhixiao Qi +4
Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long…
cs.CL2025
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
Yijiong Yu, Ziyun Dai, Zekun Wang +3
Large language models (LLMs) have demonstrated remarkable capabilities, but their success heavily relies on the quality of pretraining corpora. For Chinese LLMs, the scarcity of hi…