Reducing Overfitting in Deep Networks by Decorrelating Representations
arXiv:1511.06068
Abstract
One major challenge in training Deep Neural Networks is preventing overfitting. Many techniques such as data augmentation and novel regularizers such as Dropout have been proposed to prevent overfitting without requiring a massive amount of training data. In this work, we propose a new regularizer called DeCov which leads to significantly reduced overfitting (as indicated by the difference between train and val performance), and better generalization. Our regularizer encourages diverse or non-redundant representations in Deep Neural Networks by minimizing the cross-covariance of hidden activations. This simple intuition has been explored in a number of past works but surprisingly has never been applied as a regularizer in supervised learning. Experiments across a range of datasets and network architectures show that this loss always reduces overfitting while almost always maintaining or increasing generalization performance and often improving performance over Dropout.
12 pages, 5 figures, 5 tables, Accepted to ICLR 2016, (v4 adds acknowledgements)
Cited by in corpus (10)
- User Diverse Preference Modeling by Multimodal Attentive Metric Learning
- Transfer Learning via Contextual Invariants for One-to-Many Cross-Domain Recommendation
- Restructuring Batch Normalization to Accelerate CNN Training
- Human Pose and Path Estimation from Aerial Video using Dynamic Classifier Selection
- Continual Learning in Neural Networks
- Weakly-correlated synapses promote dimension reduction in deep neural networks
- Tuning-Free Disentanglement via Projection
- An ETF view of Dropout regularization
- Learning by Active Forgetting for Neural Networks
- Sparse-Interest Network for Sequential Recommendation