Deep regularization and direct training of the inner layers of Neural Networks with Kernel Flows
arXiv:2002.08335 · doi:10.1016/j.physd.2021.132952
Abstract
We introduce a new regularization method for Artificial Neural Networks (ANNs) based on Kernel Flows (KFs). KFs were introduced as a method for kernel selection in regression/kriging based on the minimization of the loss of accuracy incurred by halving the number of interpolation points in random batches of the dataset. Writing for the functional representation of compositional structure of the ANN, the inner layers outputs define a hierarchy of feature maps and kernels . When combined with a batch of the dataset these kernels produce KF losses (the regression error incurred by using a random half of the batch to predict the other half) depending on parameters of inner layers (and ). The proposed method simply consists in aggregating a subset of these KF losses with a classical output loss. We test the proposed method on CNNs and WRNs without alteration of structure nor output classifier and report reduced test errors, decreased generalization gaps, and increased robustness to distribution shift without significant increase in computational complexity. We suspect that these results might be explained by the fact that while conventional training only employs a linear functional (a generalized moment) of the empirical distribution defined by the dataset and can be prone to trapping in the Neural Tangent Kernel regime (under over-parameterizations), the proposed loss function (defined as a nonlinear functional of the empirical distribution) effectively trains the underlying kernel defined by the CNN beyond regressing the data with that kernel.
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Do ImageNet Classifiers Generalize to ImageNet?
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- Artificial Intelligence and Statistics
- Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
- Do CIFAR-10 Classifiers Generalize to CIFAR-10?
- Kernel Flows: from learning kernels from data into the abyss
- Cold Case: The Lost MNIST Digits
- A Kernel Perspective for Regularizing Deep Neural Networks
- Consistency of Empirical Bayes And Kernel Flow For Hierarchical Parameter Estimation
- Data-driven geophysical forecasting: Simple, low-cost, and accurate baselines with kernel methods
Cited by in corpus (6)
- SpinalNet: Deep Neural Network with Gradual Input
- Learning dynamical systems from data: a simple cross-validation perspective
- Learning "best" kernels from data in Gaussian process regression. With application to aerodynamics
- Consistency of Empirical Bayes And Kernel Flow For Hierarchical Parameter Estimation
- Deep Learning with Kernel Flow Regularization for Time Series Forecasting
- Data-driven geophysical forecasting: Simple, low-cost, and accurate baselines with kernel methods