6 papers
Ascend to Science: Exploration of AI Chips for Scientific Computing
Weicheng Xue, Kai Yang, Yongxiang Liu +5
The rapid rise of AI-oriented accelerators has reshaped compute systems around low-precision tensor engines, raising a practical question for the HPC community: under what conditio…
PowerStep: Memory-Efficient Adaptive Optimization via -Norm Steepest Descent
Yao Lu, Dengdong Fan, Shixun Zhang +1
Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of…
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
Weicheng Xue, Baisong Xu, Kai Yang +4
Modern AI accelerators provide high-throughput low-precision matrix engines, but often lack efficient support for FP32 GEMM. This paper presents SGEMM-cube, an FP32-accuracy GEMM a…
SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators
Xianglin Liu, Kai Yang, Fanli Zhou +9
The rapid advancement of deep learning is reshaping the hardware design landscape toward AI tasks, posing fundamental challenges for HPC workloads such as atomistic simulation. Her…
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
Yao Lu, Dengdong Fan, Jianzheng Nie +4
We present PCL-Reasoner-V1.5, a 32-billion-parameter large language model (LLM) for mathematical reasoning. The model is built upon Qwen2.5-32B and refined via supervised fine-tuni…
Revealing Nanostructures in High-Entropy Alloys via Machine-Learning Accelerated Scalable Monte Carlo Simulation
Xianglin Liu, Kai Yang, Yongxiang Liu +5
The computational cost of traditional first-principles method quickly becomes prohibitively expensive as the number of atoms increases. This challenge is further amplified by the n…