66 citations · 137 across the 11 of their papers we have counts for
7 papers · 1 filter
Harnessing Deep Learning and HPC Kernels via High-Level Loop and Tensor Abstractions on CPU Architectures
Evangelos Georganas, Dhiraj Kalamkar, Kirill Voronin +5
During the past decade, Deep Learning (DL) algorithms, programming systems and hardware have converged with the High Performance Computing (HPC) counterparts. Nevertheless, the pro…
FPGA-based AI Smart NICs for Scalable Distributed AI Training Systems
Rui Ma, Evangelos Georganas, Alexander Heinecke +2
Rapid advances in artificial intelligence (AI) technology have led to significant accuracy improvements in a myriad of application domains at the cost of larger and more compute-in…
Next-Generation Local Time Stepping for the ADER-DG Finite Element Method
Alexander Breuer, Alexander Heinecke
High-frequency ground motion simulations pose a grand challenge in computational seismology. Two main factors drive this challenge. First, to account for higher frequencies, we hav…
PolyDL: Polyhedral Optimizations for Creation of High Performance DL primitives
Sanket Tavarageri, Alexander Heinecke, Sasikanth Avancha +3
Deep Neural Networks (DNNs) have revolutionized many aspects of our lives. The use of DNNs is becoming ubiquitous including in softwares for image recognition, speech recognition,…
Optimizing Deep Learning Recommender Systems' Training On CPU Cluster Architectures
Dhiraj Kalamkar, Evangelos Georganas, Sudarshan Srinivasan +3
During the last two years, the goal of many researchers has been to squeeze the last bit of performance out of HPC system for AI tasks. Often this discussion is held in the context…
ISA Mapper: A Compute and Hardware Agnostic Deep Learning Compiler
Matthew Sotoudeh, Anand Venkat, Michael Anderson +3
Domain specific accelerators present new challenges and opportunities for code generation onto novel instruction sets, communication fabrics, and memory architectures. In this pape…