8 citations · 10 across the 3 of their papers we have counts for
4 papers
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
Hang Liu, Junjie Li, Yinzhi Wang
This study explores the use of automatic BLAS offloading and INT8-based emulation for accelerating traditional HPC workloads on modern GPU architectures. Through the use of low-bit…
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
Junjie Li, Yinzhi Wang, Xiao Liang +1
Porting codes to GPU often requires major efforts. While several tools exist for automatically offload numerical libraries such as BLAS and LAPACK, they often prove impractical due…
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
Bingbing Li, Geng Yuan, Zigeng Wang +6
Resistive Random Access Memory (ReRAM) has emerged as a promising platform for deep neural networks (DNNs) due to its support for parallel in-situ matrix-vector multiplication. How…
DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies
Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang +89
In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scient…