Failures of Gradient-Based Deep Learning
arXiv:1703.07950
Abstract
In recent years, Deep Learning has become the go-to solution for a broad range of applications, often outperforming state-of-the-art. However, it is important, for both theoreticians and practitioners, to gain a deeper understanding of the difficulties and limitations associated with common approaches and algorithms. We describe four types of simple problems, for which the gradient-based algorithms commonly used in deep learning either fail or suffer from significant difficulties. We illustrate the failures through practical experiments, and provide theoretical insights explaining their source, and how they might be remedied.
References in corpus (3)
Cited by in corpus (30)
- Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
- Deep Learning is Robust to Massive Label Noise
- Fisher-Rao Metric, Geometry, and Complexity of Neural Networks
- Super-Resolution via Deep Learning
- Recent advances for quantum classifiers
- The Pitfalls of Simplicity Bias in Neural Networks
- On the loss landscape of a class of deep neural networks with no bad local valleys
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Probabilistic Modeling with Matrix Product States
- Product-based Neural Networks for User Response Prediction over Multi-field Categorical Data
- Stable Tensor Neural Networks for Rapid Deep Learning
- Deep Learning as a Mixed Convex-Combinatorial Optimization Problem
- Is Deeper Better only when Shallow is Good?
- The Landscape of Deep Learning Algorithms
- Computational Separation Between Convolutional and Fully-Connected Networks
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Weight Sharing is Crucial to Succesful Optimization
- Closing the gap towards end-to-end autonomous vehicle system
- On the Blindspots of Convolutional Networks
- Disturbance Decoupling for Gradient-based Multi-Agent Learning with Quadratic Costs
- Hierarchically Compositional Tasks and Deep Convolutional Networks
- When Hardness of Approximation Meets Hardness of Learning
- A New Benchmark and Progress Toward Improved Weakly Supervised Learning
- Notes on stable learning with piecewise-linear basis functions
- Bridging Cognitive Programs and Machine Learning
- Achieving Adversarial Robustness Requires An Active Teacher
- Adaptive norms for deep learning with regularized Newton methods
- Challenging Images For Minds and Machines
- Comparison of Deep Neural Networks and Deep Hierarchical Models for Spatio-Temporal Data