collaborators

7 papers

cs.CL2026

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Dongfang Li, Xiaodong Luo, Ruoyu Sun +64

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…

cs.LG2026

A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks

Tian Ding, Dawei Li, Ruoyu Sun

We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions. We focus on the phenom…

math.OC2025

On Representing Convex Quadratically Constrained Quadratic Programs via Graph Neural Networks

Chenyang Wu, Qian Chen, Akang Wang +4

Convex quadratically constrained quadratic programs (QCQPs) involve finding a solution within a convex feasible region defined by quadratic constraints while minimizing a convex qu…

cs.LG2025

MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning

Yupeng Chen, Senmiao Wang, Yushun Zhang +5

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tu…

cs.LG2025

Learning to Gridize: Segment Physical World by Wireless Communication Channel

Juntao Wang, Feng Yin, Tian Ding +3

Gridization, the process of partitioning space into grids where users share similar channel characteristics, serves as a fundamental prerequisite for efficient large-scale network…

cs.LG2025

Exploring and Improving Initialization for Deep Graph Neural Networks: A Signal Propagation Perspective

Senmiao Wang, Yupeng Chen, Yushun Zhang +2

Graph Neural Networks (GNNs) often suffer from performance degradation as the network depth increases. This paper addresses this issue by introducing initialization methods that en…