Tensorizing Neural Networks
arXiv:1509.06569
Abstract
Deep neural networks currently demonstrate state-of-the-art performance in several domains. At the same time, models of this class are very demanding in terms of computational resources. In particular, a large amount of memory is required by commonly used fully-connected layers, making it hard to use the models on low-end devices and stopping the further increase of the model size. In this paper we convert the dense weight matrices of the fully-connected layers to the Tensor Train format such that the number of parameters is reduced by a huge factor and at the same time the expressive power of the layer is preserved. In particular, for the Very Deep VGG networks we report the compression factor of the dense weight matrix of a fully-connected layer up to 200000 times leading to the compression factor of the whole network up to 7 times.
References in corpus (4)
Cited by in corpus (142)
- The ITensor Software Library for Tensor Network Calculations
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- Convolutional Neural Networks using Logarithmic Data Representation
- Variational Dropout Sparsifies Deep Neural Networks
- Recent Advances in Convolutional Neural Networks
- Tensor Ring Decomposition
- Unsupervised Generative Modeling Using Matrix Product States
- Differentiable Programming Tensor Networks
- Tensor Networks for Dimensionality Reduction and Large-Scale Optimizations. Part 2 Applications and Future Perspectives
- TensorLy: Tensor Learning in Python
- Lecture Notes of Tensor Network Contractions
- Recurrent Neural Networks: An Embedded Computing Perspective
- PerforatedCNNs: Acceleration through Elimination of Redundant Convolutions
- Compression of Deep Convolutional Neural Networks for Fast and Low Power Mobile Applications
- Compressing Recurrent Neural Network with Tensor Train
- Review: Deep Learning in Electron Microscopy
- Ultimate tensorization: compressing convolutional and FC layers alike
- Deep Model Compression: Distilling Knowledge from Noisy Teachers
- Supervised Learning with Quantum-Inspired Tensor Networks
- The Power of Sparsity in Convolutional Neural Networks
- Loss-aware Weight Quantization of Deep Networks
- Local approximate Gaussian process regression for data-driven constitutive laws: Development and comparison with neural networks
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- Tensor Regression Networks
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Gauge fixing, canonical forms and optimal truncations in tensor networks with closed loops
- Long-term Forecasting using Higher Order Tensor RNNs
- Supervised Learning with Projected Entangled Pair States
- Structured Probabilistic Pruning for Convolutional Neural Network Acceleration
- Tensor-Train Recurrent Neural Networks for Video Classification
- Tensorized Embedding Layers for Efficient Model Compression
- Tensor Computation: A New Framework for High-Dimensional Problems in EDA
- Tensor Train Neighborhood Preserving Embedding
- From Hashing to CNNs: Training BinaryWeight Networks via Hashing
- Improving Efficiency in Convolutional Neural Network with Multilinear Filters
- Tensor Train decomposition on TensorFlow (T3F)
- Gradient-based Optimization for Regression in the Functional Tensor-Train Format
- Dynamic Graph Convolutional Networks Using the Tensor M-Product
- Neural Networks Compression for Language Modeling
- Accurate and Compact Convolutional Neural Networks with Trained Binarization
- Tensor-based algorithms for image classification
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Kronecker CP Decomposition with Fast Multiplication for Compressing RNNs
- Wide Compression: Tensor Ring Nets
- Model compression as constrained optimization, with application to neural nets. Part I: general framework
- Compressing 3DCNNs Based on Tensor Train Decomposition
- A literature survey of matrix methods for data science
- On Tensor Train Rank Minimization: Statistical Efficiency and Scalable Algorithm
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Compression and Interpretability of Deep Neural Networks via Tucker Tensor Layer: From First Principles to Tensor Valued Back-Propagation
- Resource-Efficient Neural Networks for Embedded Systems
- Stable Tensor Neural Networks for Rapid Deep Learning
- ACDC: A Structured Efficient Linear Layer
- A Dimensionality Reduction Approach for Convolutional Neural Networks
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Fast ConvNets Using Group-wise Brain Damage
- Tensorial Neural Networks: Generalization of Neural Networks and Application to Model Compression
- Learning Mixtures of Separable Dictionaries for Tensor Data: Analysis and Algorithms
- Learning Compact Recurrent Neural Networks with Block-Term Tensor Decomposition
- Nearest-Neighbor Interaction Systems in the Tensor-Train Format
- Tensor Regression Networks with various Low-Rank Tensor Approximations
- Higher-dimension Tensor Completion via Low-rank Tensor Ring Decomposition
- Multi-variable LSTM neural network for autoregressive exogenous model
- Tensor Ring Decomposition with Rank Minimization on Latent Space: An Efficient Approach for Tensor Completion
- Comprehensive SNN Compression Using ADMM Optimization and Activity Regularization
- BT-Nets: Simplifying Deep Neural Networks via Block Term Decomposition
- High-dimension Tensor Completion via Gradient-based Optimization Under Tensor-train Format
- Training Sparse Neural Networks
- Compositional Hierarchical Tensor Factorization: Representing Hierarchical Intrinsic and Extrinsic Causal Factors
- QTN-VQC: An End-to-End Learning framework for Quantum Neural Networks
- Sparse Weight Activation Training
- Tensor Contraction Layers for Parsimonious Deep Nets
- Efficient Structure-preserving Support Tensor Train Machine
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Single-Shot Matrix-Matrix Multiplication Optical Tensor Processor for Deep Learning
- The trouble with tensor ring decompositions
- Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU Platform
- Structured Pruning for Efficient ConvNets via Incremental Regularization
- PANDA: Facilitating Usable AI Development
- Adaptive Learning of Tensor Network Structures
- Tensor Decomposition for Compressing Recurrent Neural Network
- CNN Acceleration by Low-rank Approximation with Quantized Factors
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- Tensor Networks for Probabilistic Sequence Modeling
- Learning by Sampling and Compressing: Efficient Graph Representation Learning with Extremely Limited Annotations
- Tensor-Train Long Short-Term Memory for Monaural Speech Enhancement
- Walsh-Hadamard Variational Inference for Bayesian Deep Learning
- Tensorized Random Projections
- Cross-Layer Distillation with Semantic Calibration
- Efficient Low Rank Tensor Ring Completion
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition
- Understanding and Training Deep Diagonal Circulant Neural Networks
- Dynamic Sparse Graph for Efficient Deep Learning
- FALCON: Lightweight and Accurate Convolution
- Deep Learning Techniques for Compressive Sensing-Based Reconstruction and Inference -- A Ubiquitous Systems Perspective
- Compressing LSTM Networks by Matrix Product Operators
- Quantum Tensor Networks, Stochastic Processes, and Weighted Automata
- Multi-Graph Tensor Networks
- Active Subspace of Neural Networks: Structural Analysis and Universal Attacks
- Computing low-rank approximations of large-scale matrices with the Tensor Network randomized SVD
- Concatenated image completion via tensor augmentation and completion
- Provably efficient neural network representation for image classification
- Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
- A Support Tensor Train Machine
- A Unified Approximation Framework for Compressing and Accelerating Deep Neural Networks
- FeTa: A DCA Pruning Algorithm with Generalization Error Guarantees
- 3U-EdgeAI: Ultra-Low Memory Training, Ultra-Low BitwidthQuantization, and Ultra-Low Latency Acceleration
- Parallelized Tensor Train Learning of Polynomial Classifiers
- Joint Matrix Decomposition for Deep Convolutional Neural Networks Compression
- AutoHOOT: Automatic High-Order Optimization for Tensors
- A Fully Tensorized Recurrent Neural Network
- Efficient and Robust Machine Learning for Real-World Systems
- PASTA: A Parallel Sparse Tensor Algorithm Benchmark Suite
- Fast calculation of correlations in recognition systems
- Tensor Network for Supervised Learning at Finite Temperature
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Optimizing Tensor Train Decomposition in DNNs for RISC-V Architectures Using Design Space Exploration and Compiler Optimizations
- Complexity for deep neural networks and other characteristics of deep feature representations
- Tensor-generated fractals: Using tensor decompositions for creating self-similar patterns
- The Bayesian Method of Tensor Networks
- A Compact Network Learning Model for Distribution Regression
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- Spectral Tensor Train Parameterization of Deep Learning Layers
- Tensor Methods for Generating Compact Uncertainty Quantification and Deep Learning Models
- Compact Autoregressive Network
- Improving Word Embedding Factorization for Compression Using Distilled Nonlinear Neural Decomposition
- Bayesian Sparsification Methods for Deep Complex-valued Networks
- Depth Enables Long-Term Memory for Recurrent Neural Networks
- Neural network compression via learnable wavelet transforms
- Efficient Alternating Least Squares Algorithms for Low Multilinear Rank Approximation of Tensors
- MARS: Masked Automatic Ranks Selection in Tensor Decompositions
- Recurrent Graph Tensor Networks: A Low-Complexity Framework for Modelling High-Dimensional Multi-Way Sequence
- Block-term Tensor Neural Networks
- A Sampling-Based Method for Tensor Ring Decomposition
- SmartDeal: Re-Modeling Deep Network Weights for Efficient Inference and Training
- Penetrating the Fog: the Path to Efficient CNN Models
- Training Neural Machine Translation (NMT) Models using Tensor Train Decomposition on TensorFlow (T3F)
- Patch-based Medical Image Segmentation using Matrix Product State Tensor Networks
- Cyclic orthogonal convolutions for long-range integration of features
- Rademacher Random Projections with Tensor Networks
- Neural Tangent Kernel of Matrix Product States: Convergence and Applications
- Tensorization of neural networks for improved privacy and interpretability