collaborators

8 papers

cs.CL2026

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Dongfang Li, Xiaodong Luo, Ruoyu Sun +64

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…

cs.LG2026

Feature Augmentation of GNNs for ILPs: Local Uniqueness Suffices

Qingyu Han, Qian Li, Linxin Yang +3

Integer Linear Programs (ILPs) are central to real-world optimizations but notoriously difficult to solve. Learning to Optimize (L2O) has emerged as a promising paradigm, with Grap…

math.OC2025

Successive Fixing for Large-Scale SCUC Using First-Order Methods

Jinxin Xiong, Yanting Huang, Yingxiao Wang +4

Security-Constrained Unit Commitment is a fundamental optimization problem in power systems operations. The primary computational bottleneck arises from the need to solve large-sca…

cs.LG2025

QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks

Qian Chen, Linxin Yang, Akang Wang +2

The combination of linear transformations and non-linear activation functions forms the foundation of most modern deep neural networks, enabling them to approximate highly complex…

math.OC2025

Solving Quadratic Programs via Deep Unrolled Douglas-Rachford Splitting

Jinxin Xiong, Xi Gao, Linxin Yang +3

Convex quadratic programs (QPs) are fundamental to numerous applications, including finance, engineering, and energy systems. Among the various methods for solving them, the Dougla…

math.OC2025

Relax-and-Cut for Temporal SCUC Decomposition

Jinxin Xiong, Linxin Yang, Yingxiao Wang +4

The Security-Constrained Unit Commitment (SCUC) problem presents formidable computational challenges due to its combinatorial complexity, large-scale network dimensions, and numero…