Origami: A 803 GOp/s/W Convolutional Network Accelerator
arXiv:1512.04295 · doi:10.1109/TCSVT.2016.2592330
Abstract
An ever increasing number of computer vision and image/video processing challenges are being approached using deep convolutional neural networks, obtaining state-of-the-art results in object recognition and detection, semantic segmentation, action recognition, optical flow and superresolution. Hardware acceleration of these algorithms is essential to adopt these improvements in embedded and mobile computer vision systems. We present a new architecture, design and implementation as well as the first reported silicon measurements of such an accelerator, outperforming previous work in terms of power-, area- and I/O-efficiency. The manufactured device provides up to 196 GOp/s on 3.09 mm^2 of silicon in UMC 65nm technology and can achieve a power efficiency of 803 GOp/s/W. The massively reduced bandwidth requirements make it the first architecture scalable to TOp/s performance.
14 pages
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fully Convolutional Networks for Semantic Segmentation
- Deep Learning with Limited Numerical Precision
- cuDNN: Efficient Primitives for Deep Learning
- Computing the Stereo Matching Cost with a Convolutional Neural Network
- FlowNet: Learning Optical Flow with Convolutional Networks
- Training deep neural networks with low precision multiplications
- maxDNN: An Efficient Convolution Kernel for Deep Learning with Maxwell GPUs
- Compression of Deep Neural Networks on the Fly
Cited by in corpus (27)
- Hardware and Software Optimizations for Accelerating Deep Neural Networks: Survey of Current Trends, Challenges, and the Road Ahead
- Hardware Implementation of Deep Network Accelerators Towards Healthcare and Biomedical Applications
- Adaptive Extreme Edge Computing for Wearable Devices
- PULP-NN: Accelerating Quantized Neural Networks on Parallel Ultra-Low-Power RISC-V Processors
- A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
- An Electro-Photonic System for Accelerating Deep Neural Networks
- Distributed Deep Convolutional Neural Networks for the Internet-of-Things
- Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
- Taxonomy and Benchmarking of Precision-Scalable MAC Arrays Under Enhanced DNN Dataflow Representation
- FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- RPR: Random Partition Relaxation for Training; Binary and Ternary Weight Neural Networks
- FANN-on-MCU: An Open-Source Toolkit for Energy-Efficient Neural Network Inference at the Edge of the Internet of Things
- An Application-Specific VLIW Processor with Vector Instruction Set for CNN Acceleration
- Bridging the Gap Between Neural Networks and Neuromorphic Hardware with A Neural Network Compiler
- TinyCNN: A Tiny Modular CNN Accelerator for Embedded FPGA
- AnalogNet: Convolutional Neural Network Inference on Analog Focal Plane Sensor Processors
- Dataflow Aware Mapping of Convolutional Neural Networks Onto Many-Core Platforms With Network-on-Chip Interconnect
- Tuning Algorithms and Generators for Efficient Edge Inference
- Benchmarking Physical Performance of Neural Inference Circuits
- Mapping high-performance RNNs to in-memory neuromorphic chips
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Multi-Mode Inference Engine for Convolutional Neural Networks
- Tackling Variabilities in Autonomous Driving
- XpulpNN: Enabling Energy Efficient and Flexible Inference of Quantized Neural Network on RISC-V based IoT End Nodes
- FusionAccel: A General Re-configurable Deep Learning Inference Accelerator on FPGA for Convolutional Neural Networks
- Faster Convolution Inference Through Using Pre-Calculated Lookup Tables