Transport Analysis of Infinitely Deep Neural Network
arXiv:1605.02832
Abstract
We investigated the feature map inside deep neural networks (DNNs) by tracking the transport map. We are interested in the role of depth (why do DNNs perform better than shallow models?) and the interpretation of DNNs (what do intermediate layers do?) Despite the rapid development in their application, DNNs remain analytically unexplained because the hidden layers are nested and the parameters are not faithful. Inspired by the integral representation of shallow NNs, which is the continuum limit of the width, or the hidden unit number, we developed the flow representation and transport analysis of DNNs. The flow representation is the continuum limit of the depth or the hidden layer number, and it is specified by an ordinary differential equation with a vector field. We interpret an ordinary DNN as a transport map or a Euler broken line approximation of the flow. Technically speaking, a dynamical system is a natural model for the nested feature maps. In addition, it opens a new way to the coordinate-free treatment of DNNs by avoiding the redundant parametrization of DNNs. Following Wasserstein geometry, we analyze a flow in three aspects: dynamical system, continuity equation, and Wasserstein gradient flow. A key finding is that we specified a series of transport maps of the denoising autoencoder (DAE). Starting from the shallow DAE, this paper develops three topics: the transport map of the deep DAE, the equivalence between the stacked DAE and the composition of DAEs, and the development of the double continuum limit or the integral representation of the flow representation. As partial answers to the research questions, we found that deeper DAEs converge faster and the extracted features are better; in addition, a deep Gaussian DAE transports mass to decrease the Shannon entropy of the data distribution.
References in corpus (16)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- The Loss Surfaces of Multilayer Networks
- Neural Ordinary Differential Equations
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm
- The Reversible Residual Network: Backpropagation Without Storing Activations
- The loss surface of deep and wide neural networks
- No bad local minima: Data independent training error guarantees for multilayer neural networks
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Improved minimax predictive densities under Kullback--Leibler loss
- Benefits of depth in neural networks
- Implicit Regularization in Deep Learning
- An Optimal Control Approach to Deep Learning and Applications to Discrete-Weight Neural Networks
- Density in Approximation Theory
- Transportation analysis of denoising autoencoders: a novel method for analyzing deep neural networks
- Functional Gradient Boosting based on Residual Network Perception
Cited by in corpus (10)
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- You Only Propagate Once: Accelerating Adversarial Training via Maximal Principle
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Discretize-Optimize vs. Optimize-Discretize for Time-Series Regression and Continuous Normalizing Flows
- Continuous-in-Depth Neural Networks
- Feedback maximum principle for ensemble control of local continuity equations. An application to supervised machine learning
- Particle-based Energetic Variational Inference
- Stable Neural Flows
- ResNet After All? Neural ODEs and Their Numerical Solution
- No one-hidden-layer neural network can represent multivariable functions