A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
arXiv:2306.15552 · doi:10.1145/3729215
Abstract
Recent trends in deep learning (DL) have made hardware accelerators essential for various high-performance computing (HPC) applications, including image classification, computer vision, and speech recognition. This survey summarizes and classifies the most recent developments in DL accelerators, focusing on their role in meeting the performance demands of HPC applications. We explore cutting-edge approaches to DL acceleration, covering not only GPU- and TPU-based platforms but also specialized hardware such as FPGA- and ASIC-based accelerators, Neural Processing Units, open hardware RISC-V-based accelerators, and co-processors. This survey also describes accelerators leveraging emerging memory technologies and computing paradigms, including 3D-stacked Processor-In-Memory, non-volatile memories like Resistive RAM and Phase Change Memories used for in-memory computing, as well as Neuromorphic Processing Units, and Multi-Chip Module-based accelerators. Furthermore, we provide insights into emerging quantum-based accelerators and photonics. Finally, this survey categorizes the most influential architectures and technologies from recent years, offering readers a comprehensive perspective on the rapidly evolving field of deep learning acceleration.
Preprint version of our manuscript submitted to the journal @ ACM CSUR (58 pages including Appendix) on June 22nd, 2023. Major revision submitted on July 12th, 2024. Accepted for publication on March 22nd, 2025
References in corpus (28)
- Deep Learning in Neural Networks: An Overview
- Quantum Machine Learning
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Binary Neural Networks: A Survey
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- A 0.086-mm 12.7-pJ/SOP 64k-Synapse 256-Neuron Online-Learning Digital Spiking Neuromorphic Processor in 28nm CMOS
- NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
- A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference
- MorphIC: A 65-nm 738k-Synapse/mm Quad-Core Binary-Weight Digital Neuromorphic Processor with Stochastic Spike-Driven Online Learning
- MLPerf Training Benchmark
- Origami: A 803 GOp/s/W Convolutional Network Accelerator
- An Energy-Efficient FPGA-based Deconvolutional Neural Networks Accelerator for Single Image Super-Resolution
- Vega: A 10-Core SoC for IoT End-Nodes with DNN Acceleration and Cognitive Wake-Up From MRAM-Based State-Retentive Sleep Mode
- AI and ML Accelerator Survey and Trends
- A Systematic Survey of General Sparse Matrix-Matrix Multiplication
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- A "New Ara" for Vector Computing: An Open Source Highly Efficient RISC-V V 1.0 Vector Processor Design
- PERCIVAL: Open-Source Posit RISC-V Core with Quire Capability
- TinyVers: A Tiny Versatile System-on-chip with State-Retentive eMRAM for ML Inference at the Extreme Edge
- DARKSIDE: A Heterogeneous RISC-V Compute Cluster for Extreme-Edge On-Chip DNN Inference and Training
- Dustin: A 16-Cores Parallel Ultra-Low-Power Cluster with 2b-to-32b Fully Flexible Bit-Precision and Vector Lockstep Execution Mode
- SamurAI: A Versatile IoT Node With Event-Driven Wake-Up and Embedded ML Acceleration
- Spatz: A Compact Vector Processing Unit for High-Performance and Energy-Efficient Shared-L1 Clusters
- Neural-PIM: Efficient Processing-In-Memory with Neural Approximation of Peripherals
- Arrow: A RISC-V Vector Accelerator for Machine Learning Inference
- A Heterogeneous RISC-V based SoC for Secure Nano-UAV Navigation
- Adaptable Register File Organization for Vector Processors
- RedMule: A Mixed-Precision Matrix-Matrix Operation Engine for Flexible and Energy-Efficient On-Chip Linear Algebra and TinyML Training Acceleration