3 papers
cs.CL2026
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun +62
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…
hep-ex2026
Observation of in the Amplitude Analysis of
BESIII Collaboration, M. Ablikim, M. N. Achasov +716
We report the first observation of the decay in a data set corresponding to an integrated luminosity of 7.33 fb, collected in collisions by t…
cs.DC2025
High-Throughput LLM inference on Heterogeneous Clusters
Yi Xiong, Jinqi Huang, Wenjie Huang +6
Nowadays, many companies possess various types of AI accelerators, forming heterogeneous clusters. Efficiently leveraging these clusters for high-throughput large language model (L…