Regularizing Neural Networks by Penalizing Confident Output Distributions
arXiv:1701.06548
Abstract
We systematically explore regularizing neural networks by penalizing low entropy output distributions. We show that penalizing low entropy output distributions, which has been shown to improve exploration in reinforcement learning, acts as a strong regularizer in supervised learning. Furthermore, we connect a maximum entropy based confidence penalty to label smoothing through the direction of the KL divergence. We exhaustively evaluate the proposed confidence penalty and label smoothing on 6 common benchmarks: image classification (MNIST and Cifar-10), language modeling (Penn Treebank), machine translation (WMT'14 English-to-German), and speech recognition (TIMIT and WSJ). We find that both label smoothing and the confidence penalty improve state-of-the-art models across benchmarks without modifying existing hyperparameters, suggesting the wide applicability of these regularizers.
Submitted to ICLR 2017
Cited by in corpus (62)
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Fantastic Generalization Measures and Where to Find Them
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Training Neural Response Selection for Task-Oriented Dialogue Systems
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Calibration of Encoder Decoder Models for Neural Machine Translation
- Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
- Controlling Information Capacity of Binary Neural Network
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- Improved Trainable Calibration Method for Neural Networks on Medical Imaging Classification
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study
- Automatic segmentation method of pelvic floor levator hiatus in ultrasound using a self-normalising neural network
- BERT & Family Eat Word Salad: Experiments with Text Understanding
- Pre-trained Language Model Representations for Language Generation
- On Using SpecAugment for End-to-End Speech Translation
- Does label smoothing mitigate label noise?
- Learning Soft Labels via Meta Learning
- Towards Noise-resistant Object Detection with Noisy Annotations
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- The DKU-DukeECE System for the Self-Supervision Speaker Verification Task of the 2021 VoxCeleb Speaker Recognition Challenge
- Label Smoothing and Adversarial Robustness
- Confidence Adaptive Regularization for Deep Learning with Noisy Labels
- Stylistic Dialogue Generation via Information-Guided Reinforcement Learning Strategy
- Improved Adversarial Robustness via Logit Regularization Methods
- Measuring Dependence with Matrix-based Entropy Functional
- Deep Deterministic Information Bottleneck with Matrix-based Entropy Functional
- Approximating Instance-Dependent Noise via Instance-Confidence Embedding
- Applying a Pre-trained Language Model to Spanish Twitter Humor Prediction
- Improving End-To-End Modeling for Mispronunciation Detection with Effective Augmentation Mechanisms
- Instance Cross Entropy for Deep Metric Learning
- Exploiting the Full Capacity of Deep Neural Networks while Avoiding Overfitting by Targeted Sparsity Regularization
- Learning Class Unique Features in Fine-Grained Visual Classification
- DcardNet: Diabetic Retinopathy Classification at Multiple Levels Based on Structural and Angiographic Optical Coherence Tomography
- Study Group Learning: Improving Retinal Vessel Segmentation Trained with Noisy Labels
- Scalable Syntax-Aware Language Models Using Knowledge Distillation
- Bridging In- and Out-of-distribution Samples for Their Better Discriminability
- Meta-Cal: Well-controlled Post-hoc Calibration by Ranking
- Decoupled Gradient Harmonized Detector for Partial Annotation: Application to Signet Ring Cell Detection
- Adversarial Distillation for Ordered Top-k Attacks
- Introspective Learning by Distilling Knowledge from Online Self-explanation
- Maximum Entropy Regularization and Chinese Text Recognition
- Probabilistic Object Classification using CNN ML-MAP layers
- Dynamic Slimmable Network
- IIE-NLP-Eyas at SemEval-2021 Task 4: Enhancing PLM for ReCAM with Special Tokens, Re-Ranking, Siamese Encoders and Back Translation
- Friends and Foes in Learning from Noisy Labels
- ReMix: Towards Image-to-Image Translation with Limited Data
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- The Impact of Activation Sparsity on Overfitting in Convolutional Neural Networks
- Class Interference Regularization
- Deep Unsupervised Image Anomaly Detection: An Information Theoretic Framework
- Regularization via Adaptive Pairwise Label Smoothing
- Rethinking Uncertainty in Deep Learning: Whether and How it Improves Robustness
- A practical two-stage training strategy for multi-stream end-to-end speech recognition
- Classification as Decoder: Trading Flexibility for Control in Medical Dialogue
- Semi-Supervised Text Classification via Self-Pretraining
- Homography augumented momentum constrastive learning for SAR image retrieval
- Unsupervised Domain Adaptive Object Detection using Forward-Backward Cyclic Adaptation
- A Simple and Effective Approach to Automatic Post-Editing with Transfer Learning
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Identifying and Exploiting Structures for Reliable Deep Learning
- Midpoint Regularization: from High Uncertainty Training to Conservative Classification
- AugLabel: Exploiting Word Representations to Augment Labels for Face Attribute Classification