Analysis and Design of Convolutional Networks via Hierarchical Tensor Decompositions
arXiv:1705.02302
Abstract
The driving force behind convolutional networks - the most successful deep learning architecture to date, is their expressive power. Despite its wide acceptance and vast empirical evidence, formal analyses supporting this belief are scarce. The primary notions for formally reasoning about expressiveness are efficiency and inductive bias. Expressive efficiency refers to the ability of a network architecture to realize functions that require an alternative architecture to be much larger. Inductive bias refers to the prioritization of some functions over others given prior knowledge regarding a task at hand. In this paper we overview a series of works written by the authors, that through an equivalence to hierarchical tensor decompositions, analyze the expressive efficiency and inductive bias of various convolutional network architectural features (depth, width, strides and more). The results presented shed light on the demonstrated effectiveness of convolutional networks, and in addition, provide new tools for network design.
Part of the Intel Collaborative Research Institute for Computational Intelligence (ICRI-CI) Special Issue on Deep Learning Theory
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Sum-Product Networks: A New Deep Architecture
- Neural Machine Translation in Linear Time
- Deep Learning and Quantum Entanglement: Fundamental Connections with Implications to Network Design
- Boosting Dilated Convolutional Networks with Mixed Tensor Decompositions
Cited by in corpus (13)
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Implicit Regularization in Deep Learning May Not Be Explainable by Norms
- Compressing Recurrent Neural Networks Using Hierarchical Tucker Tensor Decomposition
- Computational Separation Between Convolutional and Fully-Connected Networks
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Which transformer architecture fits my data? A vocabulary bottleneck in self-attention
- Shortcut Matrix Product States and its applications
- Towards Extremely Compact RNNs for Video Recognition with Fully Decomposed Hierarchical Tucker Structure
- Implicit Regularization in Tensor Factorization
- Provably efficient neural network representation for image classification
- Complexity for deep neural networks and other characteristics of deep feature representations
- Deep Compression of Sum-Product Networks on Tensor Networks
- Learning a Self-Expressive Network for Subspace Clustering