2 papers
cs.CL2026
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
Shengrui Li, Fei Zhao, Kaiyan Zhao +6
Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such a…
cs.CL2026
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
Zihou Zhang, Zheyong Xie, Li Zhong +3
Diffusion Language Models (DLMs) have emerged as a compelling alternative to autoregressive approaches, enabling parallel text generation with competitive performance. Despite thes…