Avoiding pathologies in very deep networks
arXiv:1402.5836
Abstract
Choosing appropriate architectures and regularization strategies for deep networks is crucial to good predictive performance. To shed light on this problem, we analyze the analogous problem of constructing useful priors on compositions of functions. Specifically, we study the deep Gaussian process, a type of infinitely-wide, deep neural network. We show that in standard architectures, the representational capacity of the network tends to capture fewer degrees of freedom as the number of layers increases, retaining only a single degree of freedom in the limit. We propose an alternate network architecture which does not suffer from this pathology. We also examine deep covariance functions, obtained by composing infinitely many feature transforms. Lastly, we characterize the class of models obtained by performing dropout on Gaussian processes.
Fixed a typo regarding number of layers
References in corpus (3)
Cited by in corpus (35)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Structured and Efficient Variational Deep Learning with Matrix Gaussian Posteriors
- Variational Auto-encoded Deep Gaussian Processes
- Doubly Stochastic Variational Inference for Deep Gaussian Processes
- Nested Variational Compression in Deep Gaussian Processes
- Deep Gaussian Processes for Multi-fidelity Modeling
- All You Need is a Good Functional Prior for Bayesian Deep Learning
- Recent advances in deep learning theory
- Deep State-Space Gaussian Processes
- On the energy landscape of deep networks
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- On the expected behaviour of noise regularised deep neural networks as Gaussian processes
- Probabilistic selection of inducing points in sparse Gaussian processes
- Wide Neural Networks with Bottlenecks are Deep Gaussian Processes
- The Recurrent Neural Tangent Kernel
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective
- Compositional uncertainty in deep Gaussian processes
- Walsh-Hadamard Variational Inference for Bayesian Deep Learning
- A Survey of Techniques All Classifiers Can Learn from Deep Networks: Models, Optimizations, and Regularization
- Bayesian Alignments of Warped Multi-Output Gaussian Processes
- Beyond the Mean-Field: Structured Deep Gaussian Processes Improve the Predictive Uncertainties
- Transport Gaussian Processes for Regression
- Inter-domain Deep Gaussian Processes
- Dropout as a Regularizer of Interaction Effects
- Deep limits and cut-off phenomena for neural networks
- Compositional Modeling of Nonlinear Dynamical Systems with ODE-based Random Features
- State-space deep Gaussian processes with applications
- Deep Neural Networks as Point Estimates for Deep Gaussian Processes
- Deep Latent-Variable Kernel Learning
- Conditional Deep Gaussian Processes: multi-fidelity kernel learning
- Probabilistic Modeling for Novelty Detection with Applications to Fraud Identification
- Enhanced Recurrent Neural Tangent Kernels for Non-Time-Series Data
- Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning
- Multivariate Deep Evidential Regression
- PAC-Bayesian Bounds for Deep Gaussian Processes