17 papers
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun +64
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…
SkillSight: Calibrating Generic Content Bias for Skill Retrieval
Jinying Xiao, Bin Li, Bin Ji +9
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existi…
Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability Detection
Xin Peng, Bo Lin, Jing Wang +5
Automated vulnerability detection is crucial for enhancing software security by identifying potential flaws that attackers could exploit, thereby reducing the reliance on labor-int…
Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-Supervised Learning
Yaxin Hou, Jun Ma, Hanyang Li +3
Semi-supervised learning faces significant challenges in realistic scenarios where labeled data is scarce and unlabeled data follows unknown, arbitrary distributions. We formalize…
Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
Jinying Xiao, Bin Ji, Shasha Li +8
Large Language Models (LLMs) quantization facilitates deploying LLMs in resource-limited settings, but existing methods that combine incompatible gradient optimization and quantiza…
EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
Xiaopeng Li, Shasha Li, Xi Wang +7
Large Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted paramete…