Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices
arXiv:1603.07341 · doi:10.3389/fnins.2016.00333
Abstract
In recent years, deep neural networks (DNN) have demonstrated significant business impact in large scale analysis and classification tasks such as speech recognition, visual object detection, pattern extraction, etc. Training of large DNNs, however, is universally considered as time consuming and computationally intensive task that demands datacenter-scale computational resources recruited for many days. Here we propose a concept of resistive processing unit (RPU) devices that can potentially accelerate DNN training by orders of magnitude while using much less power. The proposed RPU device can store and update the weight values locally thus minimizing data movement during training and allowing to fully exploit the locality and the parallelism of the training algorithm. We identify the RPU device and system specifications for implementation of an accelerator chip for DNN training in a realistic CMOS-compatible technology. For large DNNs with about 1 billion weights this massively parallel RPU architecture can achieve acceleration factors of 30,000X compared to state-of-the-art microprocessors while providing power efficiency of 84,000 GigaOps/s/W. Problems that currently require days of training on a datacenter-size cluster with thousands of machines can be addressed within hours on a single RPU accelerator. A system consisted of a cluster of RPU accelerators will be able to tackle Big Data problems with trillions of parameters that is impossible to address today like, for example, natural speech recognition and translation between all world languages, real-time analytics on large streams of business and scientific data, integration and analysis of multimodal sensory data flows from massive number of IoT (Internet of Things) sensors.
19 pages, 5 figures, 2 tables
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Training and Operation of an Integrated Neuromorphic Network Based on Metal-Oxide Memristors
- Deep Learning with Limited Numerical Precision
- Integration of nanoscale memristor synapses in neuromorphic computing architectures
- Deep Image: Scaling up Image Recognition
Cited by in corpus (47)
- Neuromorphic computing with multi-memristive synapses
- Brain-inspired computing: We need a master plan
- Long short-term memory networks in memristor crossbars
- Adaptive Extreme Edge Computing for Wearable Devices
- High-Throughput In-Memory Computing for Binary Deep Neural Networks with Monolithically Integrated RRAM and 90nm CMOS
- A flexible and fast PyTorch toolkit for simulating training and inference on analog crossbar arrays
- Magnetic domain wall based synaptic and activation function generator for neuromorphic accelerators
- Mixed-precision deep learning based on computational memory
- A back-end, CMOS compatible ferroelectric Field Effect Transistor for synaptic weights
- One-step regression and classification with crosspoint resistive memory arrays
- A memristive deep belief neural network based on silicon synapses
- Using the IBM Analog In-Memory Hardware Acceleration Kit for Neural Network Training and Inference
- Adaptive Learning Rule for Hardware-based Deep Neural Networks Using Electronic Synapse Devices
- Training LSTM Networks with Resistive Cross-Point Devices
- Analog CMOS-based Resistive Processing Unit for Deep Neural Network Training
- Resource-Efficient Deep Learning: A Survey on Model-, Arithmetic-, and Implementation-Level Techniques
- Fast offset corrected in-memory training
- In-memory eigenvector computation in time O(1)
- Hardware Implementation of Spiking Neural Networks Using Time-To-First-Spike Encoding
- Nonvolatile Electrochemical Random-Access Memory Under Short Circuit
- Long Short-Term Memory Implementation Exploiting Passive RRAM Crossbar Array
- Intrinsic RESET speed limit of valence change memories
- Training large-scale ANNs on simulated resistive crossbar arrays
- Streaming Batch Eigenupdates for Hardware Neuromorphic Networks
- All-in-One Analog AI Hardware: On-Chip Training and Inference with Conductive-Metal-Oxide/HfOx ReRAM Devices
- Analytical Modelling of the Transport in Analog Filamentary Conductive-Metal-Oxide/HfOx ReRAM Devices
- Efficient Training of the Memristive Deep Belief Net Immune to Non-Idealities of the Synaptic Devices
- Design and Characterization of Superconducting Nanowire-Based Processors for Acceleration of Deep Neural Network Training
- Physical based compact model of Y-Flash memristor for neuromorphic computation
- Zero-shifting Technique for Deep Neural Network Training on Resistive Cross-point Arrays
- Efficient ConvNets for Analog Arrays
- Mixed-precision training of deep neural networks using computational memory
- Binary stochasticity enabled highly efficient neuromorphic deep learning achieves better-than-software accuracy
- Device Modeling Bias in ReRAM-based Neural Network Simulations
- Low-Rank Training of Deep Neural Networks for Emerging Memory Technology
- Update Disturbance-Resilient Analog ReRAM Crossbar Arrays for In-Memory Deep Learning Accelerators
- The Future of Computing: Bits + Neurons + Qubits
- Design Considerations for Efficient Deep Neural Networks on Processing-in-Memory Accelerators
- Recent Advances in Convolutional Neural Network Acceleration
- A temporally and spatially local spike-based backpropagation algorithm to enable training in hardware
- Design space exploration of Ferroelectric FET based Processing-in-Memory DNN Accelerator
- Hardware and software co-optimization for the initialization failure of the ReRAM based cross-bar array
- Electronegative metal dopants improve switching consistency in Al2O3 resistive switching devices
- Study of Resistive Switching Dynamics and Memory States Equilibria in Analog Filamentary Conductive-Metal-Oxide/HfOx ReRAM via Compact Modeling
- Non-Ideal Program-Time Conservation in Charge Trap Flash for Deep Learning
- SpinAPS: A High-Performance Spintronic Accelerator for Probabilistic Spiking Neural Networks
- Deep Neural Network Optimized to Resistive Memory with Nonlinear Current-Voltage Characteristics