Hardware-oriented Approximation of Convolutional Neural Networks
arXiv:1604.03168
Abstract
High computational complexity hinders the widespread usage of Convolutional Neural Networks (CNNs), especially in mobile devices. Hardware accelerators are arguably the most promising approach for reducing both execution time and power consumption. One of the most important steps in accelerator development is hardware-oriented model approximation. In this paper we present Ristretto, a model approximation framework that analyzes a given CNN with respect to numerical resolution used in representing weights and outputs of convolutional and fully connected layers. Ristretto can condense models by using fixed point arithmetic and representation instead of floating point. Moreover, Ristretto fine-tunes the resulting fixed point network. Given a maximum error tolerance of 1%, Ristretto can successfully condense CaffeNet and SqueezeNet to 8-bit. The code for Ristretto is available.
8 pages, 4 figures, Accepted as a workshop contribution at ICLR 2016. Updated comparison to other works
References in corpus (1)
Cited by in corpus (37)
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps
- Model compression via distillation and quantization
- Accelerating CNN inference on FPGAs: A Survey
- Ristretto: Hardware-Oriented Approximation of Convolutional Neural Networks
- Deep Convolutional Neural Network Inference with Floating-point Weights and Fixed-point Activations
- VIBNN: Hardware Acceleration of Bayesian Neural Networks
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- ShiftCNN: Generalized Low-Precision Architecture for Inference of Convolutional Neural Networks
- Improving Efficiency in Convolutional Neural Network with Multilinear Filters
- Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- NeuPart: Using Analytical Models to Drive Energy-Efficient Partitioning of CNN Computations on Cloud-Connected Mobile Clients
- Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
- Q-CapsNets: A Specialized Framework for Quantizing Capsule Networks
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- HCM: Hardware-Aware Complexity Metric for Neural Network Architectures
- Quantization of Deep Neural Networks for Accumulator-constrained Processors
- Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions
- A practical convolutional neural network as loop filter for intra frame
- HarDNet: A Low Memory Traffic Network
- Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA
- A Simple Method to Reduce Off-chip Memory Accesses on Convolutional Neural Networks
- Convolutional Neural Network Quantization using Generalized Gamma Distribution
- An FPGA-Accelerated Design for Deep Learning Pedestrian Detection in Self-Driving Vehicles
- A Very Compact Embedded CNN Processor Design Based on Logarithmic Computing
- DNN Feature Map Compression using Learned Representation over GF(2)
- Understanding Chat Messages for Sticker Recommendation in Messaging Apps
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Exploiting Weight Redundancy in CNNs: Beyond Pruning and Quantization
- 3U-EdgeAI: Ultra-Low Memory Training, Ultra-Low BitwidthQuantization, and Ultra-Low Latency Acceleration
- Towards Accurate and High-Speed Spiking Neuromorphic Systems with Data Quantization-Aware Deep Networks
- Face Recognition with Hybrid Efficient Convolution Algorithms on FPGAs
- IFQ-Net: Integrated Fixed-point Quantization Networks for Embedded Vision
- Method for Hybrid Precision Convolutional Neural Network Representation
- FusionAccel: A General Re-configurable Deep Learning Inference Accelerator on FPGA for Convolutional Neural Networks
- Light Multi-segment Activation for Model Compression