Convolutional neural networks with low-rank regularization
arXiv:1511.06067
Abstract
Large CNNs have delivered impressive performance in various computer vision applications. But the storage and computation requirements make it problematic for deploying these models on mobile devices. Recently, tensor decompositions have been used for speeding up CNNs. In this paper, we further develop the tensor decomposition technique. We propose a new algorithm for computing the low-rank tensor decomposition for removing the redundancy in the convolution kernels. The algorithm finds the exact global optimizer of the decomposition and is more effective than iterative methods. Based on the decomposition, we further propose a new method for training low-rank constrained CNNs from scratch. Interestingly, while achieving a significant speedup, sometimes the low-rank constrained CNNs delivers significantly better performance than their non-constrained counterparts. On the CIFAR-10 dataset, the proposed low-rank NIN model achieves accuracy (without data augmentation), which also improves upon state-of-the-art result. We evaluated the proposed method on CIFAR-10 and ILSVRC12 datasets for a variety of modern CNNs, including AlexNet, NIN, VGG and GoogleNet with success. For example, the forward time of VGG-16 is reduced by half while the performance is still comparable. Empirical success suggests that low-rank tensor decompositions can be a very useful tool for speeding up large CNNs.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Fully Convolutional Networks for Semantic Segmentation
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Learning Activation Functions to Improve Deep Neural Networks
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
Cited by in corpus (87)
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Learning Structured Sparsity in Deep Neural Networks
- A Comprehensive Survey on Model Quantization for Deep Neural Networks in Image Classification
- CirCNN: Accelerating and Compressing Deep Neural Networks Using Block-CirculantWeight Matrices
- Discrimination-aware Network Pruning for Deep Model Compression
- Compression-aware Training of Deep Networks
- Faster CNNs with Direct Sparse Convolutions and Guided Pruning
- Learning Self-Supervised Low-Rank Network for Single-Stage Weakly and Semi-Supervised Semantic Segmentation
- Deep -Means: Re-Training and Parameter Sharing with Harder Cluster Assignments for Compressing Deep Convolutions
- Tensor Regression Networks
- Soft Filter Pruning for Accelerating Deep Convolutional Neural Networks
- Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
- Towards the Limit of Network Quantization
- Combination of Hyperband and Bayesian Optimization for Hyperparameter Optimization in Deep Learning
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- Vision Transformer Pruning
- 2PFPCE: Two-Phase Filter Pruning Based on Conditional Entropy
- DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic Framework
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Improving Efficiency in Convolutional Neural Network with Multilinear Filters
- SCSP: Spectral Clustering Filter Pruning with Soft Self-adaption Manners
- Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation
- Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- ANNETTE: Accurate Neural Network Execution Time Estimation with Stacked Models
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- Learning Efficient Detector with Semi-supervised Adaptive Distillation
- Alternating Direction Method of Multipliers for Sparse Convolutional Neural Networks
- Trained Rank Pruning for Efficient Deep Neural Networks
- Knowledge Squeezed Adversarial Network Compression
- Shift-based Primitives for Efficient Convolutional Neural Networks
- RotDCF: Decomposition of Convolutional Filters for Rotation-Equivariant Deep Networks
- Fix your classifier: the marginal value of training the last weight layer
- Coordinating Filters for Faster Deep Neural Networks
- Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression
- PruneNet: Channel Pruning via Global Importance
- Constrained Deep Learning using Conditional Gradient and Applications in Computer Vision
- Rethinking Class-Discrimination Based CNN Channel Pruning
- SQuantizer: Simultaneous Learning for Both Sparse and Low-precision Neural Networks
- DeFormer: Decomposing Pre-trained Transformers for Faster Question Answering
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Sparse Neural Networks Topologies
- Structured Convolutions for Efficient Neural Network Design
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- Progressive Neural Networks for Image Classification
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- Compressed Learning of Deep Neural Networks for OpenCL-Capable Embedded Systems
- BitNet: Bit-Regularized Deep Neural Networks
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition
- Matrix and tensor decompositions for training binary neural networks
- EasyConvPooling: Random Pooling with Easy Convolution for Accelerating Training and Testing
- Predefined Sparseness in Recurrent Sequence Models
- CoDiNet: Path Distribution Modeling with Consistency and Diversity for Dynamic Routing
- Learning Recurrent Binary/Ternary Weights
- A Unified Approximation Framework for Compressing and Accelerating Deep Neural Networks
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- Initialization and Regularization of Factorized Neural Layers
- Efficient Crowd Counting via Structured Knowledge Transfer
- Low Rank Factorization for Compact Multi-Head Self-Attention
- Recent Advances in Convolutional Neural Network Acceleration
- Full-Stack Filters to Build Minimum Viable CNNs
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Joint Matrix Decomposition for Deep Convolutional Neural Networks Compression
- Blind Adversarial Pruning: Balance Accuracy, Efficiency and Robustness
- E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
- Studying the Plasticity in Deep Convolutional Neural Networks using Random Pruning
- Information-Theoretic Understanding of Population Risk Improvement with Model Compression
- Improving Network Slimming with Nonconvex Regularization
- Distributed stochastic optimization for deep learning (thesis)
- An Improving Framework of regularization for Network Compression
- MICIK: MIning Cross-Layer Inherent Similarity Knowledge for Deep Model Compression
- Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- Associative Convolutional Layers
- Distilling Knowledge From a Deep Pose Regressor Network
- Learning with Hyperspherical Uniformity
- Nonlinear Tensor Ring Network
- SlimNets: An Exploration of Deep Model Compression and Acceleration
- Adaptive Low-Rank Regularization with Damping Sequences to Restrict Lazy Weights in Deep Networks
- Radius Adaptive Convolutional Neural Network