3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.PL2021★ 3 cited
Supporting CUDA for an extended RISC-V GPU architecture
Ruobing Han, Blaise Tine, Jaewon Lee +2
With the rapid development of scientific computation, more and more researchers and developers are committed to implementing various workloads/operations on different devices. Amon…
cs.DC2019
Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
Peng Sun, Wansen Feng, Ruobing Han +2
It is important to scale out deep neural network (DNN) training for reducing model training time. The high communication overhead is one of the major performance bottlenecks for di…