29 citations · 66 across the 4 of their papers we have counts for
4 papers
Bolt: Bridging the Gap between Auto-tuners and Hardware-native Performance
Jiarong Xing, Leyuan Wang, Shang Zhang +3
Today's auto-tuners (e.g., AutoTVM, Ansor) generate efficient tensor programs by navigating a large search space to identify effective implementations, but they do so with opaque h…
UNIT: Unifying Tensorized Instruction Compilation
Jian Weng, Animesh Jain, Jie Wang +3
Because of the increasing demand for computation in DNN, researchers develope both hardware and software mechanisms to reduce the compute and memory burden. A widely adopted approa…
A Unified Optimization Approach for CNN Model Inference on Integrated GPUs
Leyuan Wang, Zhi Chen, Yizhi Liu +4
Modern deep learning applications urge to push the model inference taking place at the edge devices for multiple reasons such as achieving shorter latency, relieving the burden of…
Gunrock: GPU Graph Analytics
Yangzihao Wang, Yuechao Pan, Andrew Davidson +8
For large-scale graph analytics on the GPU, the irregularity of data access and control flow, and the complexity of programming GPUs, have presented two significant challenges to d…