3 papers
cs.LG2024
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
Haoran Lin, Xianzhi Yu, Kang Zhao +17
FlashAttention series has been widely applied in the inference of large language models (LLMs). However, FlashAttention series only supports the high-level GPU architectures, e.g.,…
cs.DC2024
Kilometer-Level Coupled Modeling Using 40 Million Cores: An Eight-Year Journey of Model Development
Xiaohui Duan, Yuxuan Li, Zhao Liu +38
With current and future leading systems adopting heterogeneous architectures, adapting existing models for heterogeneous supercomputers is of urgent need for improving model resolu…
cs.PL2023
O2ATH: An OpenMP Offloading Toolkit for the Sunway Heterogeneous Manycore Platform
Haoran Lin, Lifeng Yan, Qixin Chang +13
The next generation Sunway supercomputer employs the SW26010pro processor, which features a specialized on-chip heterogeneous architecture. Applications with significant hotspots c…