Accelerating Very Deep Convolutional Networks for Classification and Detection
arXiv:1505.06798
Abstract
This paper aims to accelerate the test-time computation of convolutional neural networks (CNNs), especially very deep CNNs that have substantially impacted the computer vision community. Unlike previous methods that are designed for approximating linear filters or linear responses, our method takes the nonlinear units into account. We develop an effective solution to the resulting nonlinear optimization problem without the need of stochastic gradient descent (SGD). More importantly, while previous methods mainly focus on optimizing one or two layers, our nonlinear method enables an asymmetric reconstruction that reduces the rapidly accumulated error when multiple (e.g., >=10) layers are approximated. For the widely used very deep VGG-16 model, our method achieves a whole-model speedup of 4x with merely a 0.3% increase of top-5 error in ImageNet classification. Our 4x accelerated VGG-16 model also shows a graceful accuracy degradation for object detection when plugged into the Fast R-CNN detector.
TPAMI, accepted. arXiv admin note: substantial text overlap with arXiv:1411.4229
References in corpus (8)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Convolutional Feature Masking for Joint Object and Stuff Segmentation
- Learning a Recurrent Visual Representation for Image Caption Generation
- Memory Bounded Deep Convolutional Networks
- Deep convolutional filter banks for texture recognition and segmentation
- Object Detection Networks on Convolutional Feature Maps
Cited by in corpus (7)
- Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications
- AutoPruner: An End-to-End Trainable Filter Pruning Method for Efficient Deep Model Inference
- Few Shot Network Compression via Cross Distillation
- Dynamic Deep Neural Networks: Optimizing Accuracy-Efficiency Trade-offs by Selective Execution
- Trained Rank Pruning for Efficient Deep Neural Networks
- Collaborative Distillation for Ultra-Resolution Universal Style Transfer
- Fully Quantized Image Super-Resolution Networks