3 papers
cs.DC2026
Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite Elements
Yinuo Wang, Lin Gan, Tianqi Mao +8
Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage…
cs.DC2026
StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence
Wenxuan Zhao, Yingfa Chen, Xu Han +7
Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency. Ho…
cs.DC2025
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
Yinuo Wang, Tianqi Mao, Lin Gan +8
Matrix-accelerated stencil computation is a hot research topic, yet its application to three-dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence…