Understanding training and generalization in deep learning by Fourier analysis
arXiv:1808.04295
Abstract
Background: It is still an open research area to theoretically understand why Deep Neural Networks (DNNs)---equipped with many more parameters than training data and trained by (stochastic) gradient-based methods---often achieve remarkably low generalization error. Contribution: We study DNN training by Fourier analysis. Our theoretical framework explains: i) DNN with (stochastic) gradient-based methods often endows low-frequency components of the target function with a higher priority during the training; ii) Small initialization leads to good generalization ability of DNN while preserving the DNN's ability to fit any function. These results are further confirmed by experiments of DNNs fitting the following datasets, that is, natural images, one-dimensional functions and MNIST dataset.
10 pages, 4 figures
References in corpus (5)
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- A Closer Look at Memorization in Deep Networks
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- Training behavior of deep neural network in frequency domain
Cited by in corpus (24)
- Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
- The Convergence Rate of Neural Networks for Learned Functions of Different Frequencies
- Review: Deep Learning in Electron Microscopy
- Towards Understanding the Spectral Bias of Deep Learning
- Training behavior of deep neural network in frequency domain
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- Multi-scale Deep Neural Networks for Solving High Dimensional PDEs
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Discovering and Explaining the Representation Bottleneck of DNNs
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- Frequency Principle in Deep Learning with General Loss Functions and Its Potential Application
- A Phase Shift Deep Neural Network for High Frequency Approximation and Wave Problems
- PhaseDNN - A Parallel Phase Shift Deep Neural Network for Adaptive Wideband Learning
- The Slow Deterioration of the Generalization Error of the Random Feature Model
- Deep frequency principle towards understanding why deeper learning is faster
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Trap of Feature Diversity in the Learning of MLPs
- Exploring The Effect of High-frequency Components in GANs Training
- Implicit Regularization in Tensor Factorization
- An Upper Limit of Decaying Rate with Respect to Frequency in Deep Neural Network
- Nonlinear Collaborative Scheme for Deep Neural Networks
- A priori generalization error for two-layer ReLU neural network through minimum norm solution
- Is the Meta-Learning Idea Able to Improve the Generalization of Deep Neural Networks on the Standard Supervised Learning?
- Diagnosing Convolutional Neural Networks using their Spectral Response