collaborators

6 papers

cs.DC2026

Ascend to Science: Exploration of AI Chips for Scientific Computing

Weicheng Xue, Kai Yang, Yongxiang Liu +5

The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditio…

cs.LG2026

PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent

Yao Lu, Dengdong Fan, Shixun Zhang +1

Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of…

cs.DC2026

SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines

Weicheng Xue, Baisong Xu, Kai Yang +4

Modern AI accelerators provide high-throughput low-precision matrix engines, but often lack efficient support for FP32 GEMM. This paper presents SGEMM-cube, an FP32-accuracy GEMM a…

physics.comp-ph2026

SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators

Xianglin Liu, Kai Yang, Fanli Zhou +9

The rapid advancement of deep learning is reshaping the hardware design landscape toward AI tasks, posing fundamental challenges for HPC workloads such as atomistic simulation. Her…

cs.LG2026

PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning

Yao Lu, Dengdong Fan, Jianzheng Nie +4

We present PCL-Reasoner-V1.5, a 32-billion-parameter large language model (LLM) for mathematical reasoning. The model is built upon Qwen2.5-32B and refined via supervised fine-tuni…

cond-mat.mtrl-sci2025

Revealing Nanostructures in High-Entropy Alloys via Machine-Learning Accelerated Scalable Monte Carlo Simulation

Xianglin Liu, Kai Yang, Yongxiang Liu +5

The computational cost of traditional first-principles method quickly becomes prohibitively expensive as the number of atoms increases. This challenge is further amplified by the n…