Speeding up Convolutional Neural Networks with Low Rank Expansions
arXiv:1405.3866
Abstract
The focus of this paper is speeding up the evaluation of convolutional neural networks. While delivering impressive results across a range of computer vision and machine learning tasks, these networks are computationally demanding, limiting their deployability. Convolutional layers generally consume the bulk of the processing time, and so in this work we present two simple schemes for drastically speeding up these layers. This is achieved by exploiting cross-channel or filter redundancy to construct a low rank basis of filters that are rank-1 in the spatial domain. Our methods are architecture agnostic, and can be easily applied to existing CPU and GPU convolutional frameworks for tuneable speedup performance. We demonstrate this with a real world network designed for scene text character recognition, showing a possible 2.5x speedup with no loss in accuracy, and 4.5x speedup with less than 1% drop in accuracy, still achieving state-of-the-art on standard benchmarks.
References in corpus (2)
Cited by in corpus (109)
- LoRA: Low-Rank Adaptation of Large Language Models
- FitNets: Hints for Thin Deep Nets
- Compressing Deep Convolutional Networks using Vector Quantization
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Pruning Filters for Efficient ConvNets
- Learning Structured Sparsity in Deep Neural Networks
- Channel Pruning for Accelerating Very Deep Neural Networks
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- Scene Text Detection via Holistic, Multi-Channel Prediction
- Memory Bounded Deep Convolutional Networks
- Focus: Querying Large Video Datasets with Low Latency and Low Cost
- The Power of Sparsity in Convolutional Neural Networks
- Interleaved Group Convolutions for Deep Neural Networks
- DyNet: Dynamic Convolution for Accelerating Convolutional Neural Networks
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Effective Quantization Methods for Recurrent Neural Networks
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- MEC: Memory-efficient Convolution for Deep Neural Network
- Compressing Convolutional Neural Networks
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Convolutional Neural Networks at Constrained Time Cost
- Learning to Prune Filters in Convolutional Neural Networks
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Knowledge Distillation Methods for Efficient Unsupervised Adaptation Across Multiple Domains
- Performance Guaranteed Network Acceleration via High-Order Residual Quantization
- Evolutionary Synthesis of Deep Neural Networks via Synaptic Cluster-driven Genetic Encoding
- Model Compression with Multi-Task Knowledge Distillation for Web-scale Question Answering System
- Knowledge Projection for Deep Neural Networks
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box Model
- Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Trained Rank Pruning for Efficient Deep Neural Networks
- Alternating Direction Method of Multipliers for Sparse Convolutional Neural Networks
- Knowledge Squeezed Adversarial Network Compression
- Search for Better Students to Learn Distilled Knowledge
- Grid Loss: Detecting Occluded Faces
- BT-Nets: Simplifying Deep Neural Networks via Block Term Decomposition
- Compressing Neural Language Models by Sparse Word Representations
- Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
- PruneNet: Channel Pruning via Global Importance
- Rethinking Class-Discrimination Based CNN Channel Pruning
- XSepConv: Extremely Separated Convolution
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- CUP: Cluster Pruning for Compressing Deep Neural Networks
- Learning Versatile Convolution Filters for Efficient Visual Recognition
- Exploring Gradient Flow Based Saliency for DNN Model Compression
- Multiscale Hierarchical Convolutional Networks
- Constrained deep neural network architecture search for IoT devices accounting hardware calibration
- Improving Neural Network Training in Low Dimensional Random Bases
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System
- Compressing complex convolutional neural network based on an improved deep compression algorithm
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Pufferfish: Communication-efficient Models At No Extra Cost
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- LPRNet: Lightweight Deep Network by Low-rank Pointwise Residual Convolution
- Binarized Convolutional Neural Networks with Separable Filters for Efficient Hardware Acceleration
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization
- Finding trainable sparse networks through Neural Tangent Transfer
- Dynamic Group Convolution for Accelerating Convolutional Neural Networks
- REPrune: Filter Pruning via Representative Election
- HGC: Hierarchical Group Convolution for Highly Efficient Neural Network
- Correlation Congruence for Knowledge Distillation
- Optimal Quantization for Batch Normalization in Neural Network Deployments and Beyond
- Principal Component Networks: Parameter Reduction Early in Training
- Adaptive Low-Rank Factorization to regularize shallow and deep neural networks
- Network Automatic Pruning: Start NAP and Take a Nap
- Parameter Efficient Deep Neural Networks with Bilinear Projections
- NeRV: Neural Representations for Videos
- Compact representations of convolutional neural networks via weight pruning and quantization
- Knowledge Distillation via Instance-level Sequence Learning
- Fast Training of Convolutional Neural Networks via Kernel Rescaling
- BasisConv: A method for compressed representation and learning in CNNs
- Deeplite Neutrino: An End-to-End Framework for Constrained Deep Learning Model Optimization
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Convolutional Hough Matching Networks for Robust and Efficient Visual Correspondence
- A Feature-map Discriminant Perspective for Pruning Deep Neural Networks
- Towards Deep Compositional Networks
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Deep Neural Network Approximation using Tensor Sketching
- E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- CircConv: A Structured Convolution with Low Complexity
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- Building Fast and Compact Convolutional Neural Networks for Offline Handwritten Chinese Character Recognition
- Mining the Weights Knowledge for Optimizing Neural Network Structures
- Semi-tensor Product-based TensorDecomposition for Neural Network Compression
- Artificial Intelligence for 5G Wireless Systems: Opportunities, Challenges, and Future Research Directions
- Radius Adaptive Convolutional Neural Network
- Adaptive Low-Rank Regularization with Damping Sequences to Restrict Lazy Weights in Deep Networks
- Knowledge Distillation By Sparse Representation Matching
- Class-Discriminative CNN Compression
- Balancing Accuracy and Latency in Multipath Neural Networks
- Magnitude and Uncertainty Pruning Criterion for Neural Networks
- Towards Efficient Tensor Decomposition-Based DNN Model Compression with Optimization Framework
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal Matrices
- Architecture Aware Latency Constrained Sparse Neural Networks
- Dep-: Improving -based Network Sparsification via Dependency Modeling
- DAC: Data-free Automatic Acceleration of Convolutional Networks
- FastSal: a Computationally Efficient Network for Visual Saliency Prediction
- Compression-aware Continual Learning using Singular Value Decomposition
- Multi-objective Evolutionary Approach for Efficient Kernel Size and Shape for CNN
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- Parallel Blockwise Knowledge Distillation for Deep Neural Network Compression
- Prune the Convolutional Neural Networks with Sparse Shrink