On the Expressive Power of Deep Learning: A Tensor Analysis
arXiv:1509.05009
Abstract
It has long been conjectured that hypotheses spaces suitable for data that is compositional in nature, such as text or images, may be more efficiently represented with deep hierarchical networks than with shallow ones. Despite the vast empirical evidence supporting this belief, theoretical justifications to date are limited. In particular, they do not account for the locality, sharing and pooling constructs of convolutional networks, the most successful deep learning architecture to date. In this work we derive a deep network architecture based on arithmetic circuits that inherently employs locality, sharing and pooling. An equivalence between the networks and hierarchical tensor factorizations is established. We show that a shallow network corresponds to CP (rank-1) decomposition, whereas a deep network corresponds to Hierarchical Tucker decomposition. Using tools from measure theory and matrix algebra, we prove that besides a negligible set, all functions that can be implemented by a deep network of polynomial size, require exponential size in order to be realized (or even approximated) by a shallow network. Since log-space computation transforms our networks into SimNets, the result applies directly to a deep learning architecture demonstrating promising empirical performance. The construction and theory developed in this paper shed new light on various practices and ideas employed by the deep learning community.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sum-Product Networks: A New Deep Architecture
- Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition
- The Power of Depth for Feedforward Neural Networks
- Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods
- Representation Benefits of Deep Feedforward Networks
- Deep Convolutional Networks are Hierarchical Kernel Machines
Cited by in corpus (40)
- Unsupervised Generative Modeling Using Matrix Product States
- Robust Large Margin Deep Neural Networks
- Tensor Networks for Dimensionality Reduction and Large-Scale Optimizations. Part 2 Applications and Future Perspectives
- Quantum Entanglement in Deep Learning Architectures
- The Power of Depth for Feedforward Neural Networks
- TensorLy: Tensor Learning in Python
- Convergence of the Deep BSDE Method for Coupled FBSDEs
- Representation Benefits of Deep Feedforward Networks
- Multi-Layer Convolutional Sparse Modeling: Pursuit and Dictionary Learning
- Deep Polynomial Neural Networks
- On the Expressive Power of Deep Neural Networks
- Two-hidden-layer Feedforward Neural Networks are Universal Approximators: A Constructive Approach
- Hybrid Tensor Decomposition in Neural Network Compression
- Convolutional Neural Nets in Chemical Engineering: Foundations, Computations, and Applications
- On the ability of neural nets to express distributions
- Tensor-based algorithms for image classification
- Fourth-order Tensors with Multidimensional Discrete Transforms
- Wide Compression: Tensor Ring Nets
- Provably scale-covariant continuous hierarchical networks based on scale-normalized differential expressions coupled in cascade
- Convolutional Rectifier Networks as Generalized Tensor Decompositions
- Mutual Information Scaling for Tensor Network Machine Learning
- Efficient Representation of Low-Dimensional Manifolds using Deep Networks
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Learning with tree tensor networks: complexity estimates and model selection
- Nearest-Neighbor Interaction Systems in the Tensor-Train Format
- Do Neural Nets Learn Statistical Laws behind Natural Language?
- Multi-Scale and Multi-Layer Contrastive Learning for Domain Generalization
- Tensor network approaches for learning non-linear dynamical laws
- Tensor Contraction Layers for Parsimonious Deep Nets
- Interpretable Convolutional Neural Networks via Feedforward Design
- CNNs Avoid Curse of Dimensionality by Learning on Patches
- Language Modeling with Reduced Densities
- From Deep to Shallow: Transformations of Deep Rectifier Networks
- A quantum inspired approach to learning dynamical laws from data -- block-sparsity and gauge-mediated weight sharing
- On the Compressive Power of Deep Rectifier Networks for High Resolution Representation of Class Boundaries
- PASTA: A Parallel Sparse Tensor Algorithm Benchmark Suite
- Efficient Deep Learning of GMMs
- How ConvNets model Non-linear Transformations
- Stochastic Function Norm Regularization of Deep Networks
- Understanding Convolutional Neural Networks with A Mathematical Model