2 citations · 3 across the 3 of their papers we have counts for
4 papers
Benchmarking GPU and TPU Performance with Graph Neural Networks
xiangyang Ju, Yunsong Wang, Daniel Murnane +3
Many artificial intelligence (AI) devices have been developed to accelerate the training and inference of neural networks models. The most common ones are the Graphics Processing U…
Porting CMS Heterogeneous Pixel Reconstruction to Kokkos
Taylor Childers, Matti J. Kortelainen, Martin Kwok +2
Programming for a diverse set of compute accelerators in addition to the CPU is a challenge. Maintaining separate source code for each architecture would require lots of effort, an…
Time-Based Roofline for Deep Learning Performance Analysis
Yunsong Wang, Charlene Yang, Steven Farrell +3
Deep learning applications are usually very compute-intensive and require a long run time for training and inference. This has been tackled by researchers from both hardware and so…
Hierarchical Roofline Performance Analysis for Deep Learning Applications
Charlene Yang, Yunsong Wang, Steven Farrell +2
This paper presents a practical methodology for collecting performance data necessary to conduct hierarchical Roofline analysis on NVIDIA GPUs. It discusses the extension of the Em…