Ultimate tensorization: compressing convolutional and FC layers alike
arXiv:1611.03214
Abstract
Convolutional neural networks excel in image recognition tasks, but this comes at the cost of high computational and memory complexity. To tackle this problem, [1] developed a tensor factorization framework to compress fully-connected layers. In this paper, we focus on compressing convolutional layers. We show that while the direct application of the tensor framework [1] to the 4-dimensional kernel of convolution does compress the layer, we can do better. We reshape the convolutional kernel into a tensor of higher order and factorize it. We combine the proposed approach with the previous work to compress both convolutional and fully-connected layers of a network and achieve 80x network compression rate with 1.1% accuracy drop on the CIFAR-10 dataset.
NIPS 2016 workshop: Learning with Tensors: Why Now and How?
References in corpus (1)
Cited by in corpus (18)
- Tensor Methods in Computer Vision and Deep Learning
- Tensor-Train Recurrent Neural Networks for Video Classification
- Hybrid Tensor Decomposition in Neural Network Compression
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Bayesian Tensorized Neural Networks with Automatic Rank Selection
- Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
- BT-Nets: Simplifying Deep Neural Networks via Block Term Decomposition
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- Tensorizing GAN with High-Order Pooling for Alzheimer's Disease Assessment
- A flexible, extensible software framework for model compression based on the LC algorithm
- Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement
- Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product Operators
- Semi-tensor Product-based TensorDecomposition for Neural Network Compression
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Towards Efficient Tensor Decomposition-Based DNN Model Compression with Optimization Framework
- Tensor Yard: One-Shot Algorithm of Hardware-Friendly Tensor-Train Decomposition for Convolutional Neural Networks
- Low-Rank+Sparse Tensor Compression for Neural Networks