18 citations · 28 across the 3 of their papers we have counts for
3 papers
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
Hongzheng Chen, Bin Fan, Alexander Collins +8
Modern GPUs feature specialized hardware units that enable high-performance, asynchronous dataflow execution. However, the conventional SIMT programming model is fundamentally misa…
Tensor Program Optimization with Probabilistic Programs
Junru Shao, Xiyou Zhou, Siyuan Feng +7
Automatic optimization for tensor programs becomes increasingly important as we deploy deep learning in various environments, and efficient optimization relies on a rich search spa…
Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
Animesh Jain, Shoubhik Bhattacharya, Masahiro Masuda +2
A growing number of applications implement predictive functions using deep learning models, which require heavy use of compute and memory. One popular technique for increasing reso…