7 citations · 15 across the 12 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun +62
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…
cs.CL2025
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives
Yajiao Liu, Congliang Chen, Junchi Yang +1
Training large language models with data collected from various domains can improve their performance on downstream tasks. However, given a fixed training budget, the sampling prop…