From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Scaling Domain Data Repetition in LLM Pretraining
Jingwei Li, Xinran Gu, Rui Dai +5
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)). However, high-quality doma…
cs.LG2026
Explaining Data Mixing Scaling Laws
Rui Dai, Shuran Zheng
The paper proposes a theoretical framework that explains how mixing data from multiple domains affects model loss, identifying capacity competition and noise reduction as key mecha…
cs.LG2025
Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
Rui Dai, Sile Hu, Xu Shen +3
Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates tha…