3 papers
cs.AI2025
Scaling up Masked Diffusion Models on Text
Shen Nie, Fengqi Zhu, Chao Du +5
Masked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language underst…
cs.CL2025
RegMix: Data Mixture as Regression for Language Model Pre-training
Qian Liu, Xiaosen Zheng, Niklas Muennighoff +5
The data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatica…
cs.CL2024
SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages
Jia Guo, Longxu Dou, Guangtao Zeng +3
In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass e…