4 papers
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
Hang Liu, Junjie Li, Yinzhi Wang +2
This study explores the use of INT8-based emulation for accelerating traditional FP64-based HPC workloads on modern GPU architectures. Through SCILIB-Accel automatic BLAS offload t…
Performant Automatic BLAS Offloading on Unified Memory Architecture with OpenMP First-Touch Style Data Movement
Junjie Li
BLAS is a fundamental building block of advanced linear algebra libraries and many modern scientific computing applications. GPUs are known for their strong arithmetic computing ca…
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
Junjie Li, Yinzhi Wang, Xiao Liang +1
Porting codes to GPU often requires major efforts. While several tools exist for automatically offload numerical libraries such as BLAS and LAPACK, they often prove impractical due…
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
Hang Liu, Junjie Li, Yinzhi Wang
This study explores the use of automatic BLAS offloading and INT8-based emulation for accelerating traditional HPC workloads on modern GPU architectures. Through the use of low-bit…