activity
20182024
most citedAN5D: Automated Stencil Framework for High-Degree Temporal Blocking on GPUs

60 citations · 78 across the 7 of their papers we have counts for

collaborators

11 papers

cs.DC20222 cited

Preparing for the Future -- Rethinking Proxy Apps

Satoshi Matsuoka, Jens Domke, Mohamed Wahib +5

A considerable amount of research and engineering went into designing proxy applications, which represent common high-performance computing workloads, to co-design and evaluate the…

cs.LG20213 cited

MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems

Steven Farrell, Murali Emani, Jacob Balma +40

Scientific communities are increasingly adopting machine learning and deep learning models in their applications to accelerate scientific insights. High performance computing syste…

cs.DC2021

Performance Portable Back-projection Algorithms on CPUs: Agnostic Data Locality and Vectorization Optimizations

Peng Chen, Mohamed Wahib, Xiao Wang +4

Computed Tomography (CT) is a key 3D imaging technology that fundamentally relies on the compute-intense back-projection operation to generate 3D volumes. GPUs are typically used f…

cs.DC202111 cited

An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks

Albert Njoroge Kahira, Truong Thao Nguyen, Leonardo Bautista Gomez +3

Deep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using…

cs.DC2020

Matrix Engines for High Performance Computing:A Paragon of Performance or Grasping at Straws?

Jens Domke, Emil Vatai, Aleksandr Drozd +8

Matrix engines or units, in different forms and affinities, are becoming a reality in modern processors; CPUs and otherwise. The current and dominant algorithmic approach to Deep L…

cs.DC20201 cited

Scaling Distributed Deep Learning Workloads beyond the Memory Capacity with KARMA

Mohamed Wahib, Haoyu Zhang, Truong Thao Nguyen +5

The dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a via…