The Expressive Power of Neural Networks: A View from the Width
arXiv:1709.02540
Abstract
The expressive power of neural networks is important for understanding deep learning. Most existing works consider this problem from the view of the depth of a network. In this paper, we study how width affects the expressiveness of neural networks. Classical results state that depth-bounded (e.g. depth-) networks with suitable activation functions are universal approximators. We show a universal approximation theorem for width-bounded ReLU networks: width- ReLU networks, where is the input dimension, are universal approximators. Moreover, except for a measure zero set, all functions cannot be approximated by width- ReLU networks, which exhibits a phase transition. Several recent works demonstrate the benefits of depth by proving the depth-efficiency of neural networks. That is, there are classes of deep networks which cannot be realized by any shallow network whose size is no more than an exponential bound. Here we pose the dual question on the width-efficiency of ReLU networks: Are there wide networks that cannot be realized by narrow networks whose size is not substantially larger? We show that there exist classes of wide networks which cannot be realized by any narrow network whose depth is no more than a polynomial bound. On the other hand, we demonstrate by extensive experiments that narrow networks whose size exceed the polynomial bound by a constant factor can approximate wide and shallow network with high accuracy. Our results provide more comprehensive evidence that depth is more effective than width for the expressiveness of ReLU networks.
accepted by NIPS 2017 ( with some typos fixed)
References in corpus (1)
Cited by in corpus (23)
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Hyperbolic compactification of M-theory and de Sitter quantum gravity
- Elvet -- a neural network-based differential equation and variational problem solver
- Deep Learning in Wide-field Surveys: Fast Analysis of Strong Lenses in Ground-based Cosmic Experiments
- Deep Nonparametric Regression on Approximate Manifolds: Non-Asymptotic Error Bounds with Polynomial Prefactors
- Model reduction in acoustic inversion by artificial neural network
- On approximating with neural networks
- A Machine Learning Approach to Correcting Atmospheric Seeing in Solar Flare Observations
- Robust Nonparametric Regression with Deep Neural Networks
- Deep KKL: Data-driven Output Prediction for Non-Linear Systems
- On Universal Approximation by Neural Networks with Uniform Guarantees on Approximation of Infinite Dimensional Maps
- Neural Quantum State Study of Fracton Models
- Dynamic Neural Diversification: Path to Computationally Sustainable Neural Networks
- A Bregman Learning Framework for Sparse Neural Networks
- Learning to Embed Categorical Features without Embedding Tables for Recommendation
- Non-asymptotic Excess Risk Bounds for Classification with Deep Convolutional Neural Networks
- Scalable Unidirectional Pareto Optimality for Multi-Task Learning with Constraints
- Fourier Neural Networks for Function Approximation
- Scalable Partial Explainability in Neural Networks via Flexible Activation Functions
- MixMix: All You Need for Data-Free Compression Are Feature and Data Mixing
- Parallel frequency function-deep neural network for efficient complex broadband signal approximation
- Bayesian inference of numerical modeling-based morphodynamics: Application to a dam-break over a mobile bed experiment