1 citations · 2 across the 2 of their papers we have counts for
3 papers
cs.DC2020
Matrix Engines for High Performance Computing:A Paragon of Performance or Grasping at Straws?
Jens Domke, Emil Vatai, Aleksandr Drozd +8
Matrix engines or units, in different forms and affinities, are becoming a reality in modern processors; CPUs and otherwise. The current and dominant algorithmic approach to Deep L…
cs.DC2020★ 1 cited
Scaling Distributed Deep Learning Workloads beyond the Memory Capacity with KARMA
Mohamed Wahib, Haoyu Zhang, Truong Thao Nguyen +5
The dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a via…
cs.DC2020★ 1 cited
A Study of Single and Multi-device Synchronization Methods in Nvidia GPUs
Lingqi Zhang, Mohamed Wahib, Haoyu Zhang +1
GPUs are playing an increasingly important role in general-purpose computing. Many algorithms require synchronizations at different levels of granularity in a single GPU. Additiona…