2 papers
cs.LG2026
CausalMix: Data Mixture as Causal Inference for Language Model Training
Zinan Tang, Yukun Zhang, Shaomian Zheng +6
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely o…
cs.CL2026
Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages
Haolin Li, Haipeng Zhang, Mang Li +4
Large language models (LLMs) continue to struggle with low-resource languages, primarily due to limited training data, translation noise, and unstable cross-lingual alignment. To a…