collaborators

8 papers

cs.CL2026

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Dongfang Li, Xiaodong Luo, Ruoyu Sun +64

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…

cs.LG2026

Adam Converges Without Any Modification On Update Rules

Yushun Zhang, Bingran Li, Congliang Chen +2

Adam is the default algorithm for training neural networks, including large language models (LLMs). However, \citet{reddi2019convergence} provided an example that Adam diverges, ra…

cs.LG2025

MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning

Yupeng Chen, Senmiao Wang, Yushun Zhang +5

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tu…

cs.LG2025

Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation

Ziniu Li, Congliang Chen, Tianyun Yang +5

Large Language Models (LLMs) can self-improve through reinforcement learning, where they generate trajectories to explore and discover better solutions. However, this exploration p…

cs.LG2025

Towards Quantifying the Hessian Structure of Neural Networks

Zhaorui Dong, Yushun Zhang, Jianfeng Yao +1

Empirical studies reported that the Hessian matrix of neural networks (NNs) exhibits a near-block-diagonal structure, yet its theoretical foundation remains unclear. In this work,…

cs.CL2025

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives

Yajiao Liu, Congliang Chen, Junchi Yang +1

Training large language models with data collected from various domains can improve their performance on downstream tasks. However, given a fixed training budget, the sampling prop…