collaborators

5 papers

cs.DC2026

Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite Elements

Yinuo Wang, Lin Gan, Tianqi Mao +8

Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage…

cs.DC2026

High-Order Spectral Element Methods for Wave Propagation on ARM Multicore CPU with SME: Optimizations and Implications

Yinuo Wang, Lin Gan, Tianqi Mao +5

Wave propagation based on the spectral element method (SEM) is a representative HPC workload, but existing SEM implementations are not well matched to emerging ARM multicore CPUs w…

cs.DC2025

pdGRASS: A Fast Parallel Density-Aware Algorithm for Graph Spectral Sparsification

Tiancheng Zhao, Zekun Yin, Huihai An +4

Graph Spectral Sparsification (GSS) identifies an ultra-sparse subgraph, or sparsifier, whose Laplacian matrix closely approximates the spectral properties of the original graph, e…

cs.DC2025

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference

Haoran Lin, Xianzhi Yu, Kang Zhao +7

Current inference systems for Mixture-of-Experts (MoE) models primarily employ static parallelization strategies. However, these static approaches cannot consistently achieve optim…

cs.DC2025

MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit

Yinuo Wang, Tianqi Mao, Lin Gan +8

Matrix-accelerated stencil computation is a hot research topic, yet its application to three-dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence…