Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels
arXiv:1805.07836
Abstract
Deep neural networks (DNNs) have achieved tremendous success in a variety of applications across many disciplines. Yet, their superior performance comes with the expensive cost of requiring correctly annotated large-scale datasets. Moreover, due to DNNs' rich capacity, errors in training labels can hamper performance. To combat this problem, mean absolute error (MAE) has recently been proposed as a noise-robust alternative to the commonly-used categorical cross entropy (CCE) loss. However, as we show in this paper, MAE can perform poorly with DNNs and challenging datasets. Here, we present a theoretically grounded set of noise-robust loss functions that can be seen as a generalization of MAE and CCE. Proposed loss functions can be readily applied with any existing DNN architecture and algorithm, while yielding good performance in a wide range of noisy label scenarios. We report results from experiments conducted with CIFAR-10, CIFAR-100 and FASHION-MNIST datasets and synthetically generated noisy labels.
32nd Conference on Neural Information Processing Systems (NeurIPS 2018)
References in corpus (14)
- Understanding deep learning requires rethinking generalization
- Learning to Reweight Examples for Robust Deep Learning
- Learning Deconvolution Network for Semantic Segmentation
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
- Training Convolutional Networks with Noisy Labels
- A Closer Look at Memorization in Deep Networks
- Masking: A New Perspective of Noisy Supervision
- Robust Loss Functions under Label Noise for Deep Neural Networks
- Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
- Learning From Noisy Singly-labeled Data
- Joint Optimization Framework for Learning with Noisy Labels
- Maximum L-likelihood estimation
- Iterative Learning with Open-set Noisy Labels
Cited by in corpus (66)
- Deep Learning for Anomaly Detection in Log Data: A Survey
- Supervised Contrastive Learning
- Active label cleaning for improved dataset quality under resource constraints
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- WaferSegClassNet -- A Light-weight Network for Classification and Segmentation of Semiconductor Wafer Defects
- Contrast to Divide: Self-Supervised Pre-Training for Learning with Noisy Labels
- Complementary Pseudo Labels For Unsupervised Domain Adaptation On Person Re-identification
- In-situ crack and keyhole pore detection in laser directed energy deposition through acoustic signal and deep learning
- OSLNet: Deep Small-Sample Classification with an Orthogonal Softmax Layer
- Low-Quality Training Data Only? A Robust Framework for Detecting Encrypted Malicious Network Traffic
- EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
- On the Effects of Different Types of Label Noise in Multi-Label Remote Sensing Image Classification
- Deep Ice Layer Tracking and Thickness Estimation using Fully Convolutional Networks
- Learning from Noisy Labels for Entity-Centric Information Extraction
- Boosting Facial Expression Recognition by A Semi-Supervised Progressive Teacher
- DSAM: A Deep Learning Framework for Analyzing Temporal and Spatial Dynamics in Brain Networks
- Regularizing Neural Network Training via Identity-wise Discriminative Feature Suppression
- Bag of Tricks for Node Classification with Graph Neural Networks
- Drastic Circuit Depth Reductions with Preserved Adversarial Robustness by Approximate Encoding for Quantum Machine Learning
- Robust Local Preserving and Global Aligning Network for Adversarial Domain Adaptation
- A comparative study of 2D image segmentation algorithms for traumatic brain lesions using CT data from the ProTECTIII multicenter clinical trial
- Privacy-Preserving Ensemble Infused Enhanced Deep Neural Network Framework for Edge Cloud Convergence
- On Using Machine Learning to Identify Knowledge in API Reference Documentation
- MAG-Net: Multi-task attention guided network for brain tumor segmentation and classification
- MetaLabelNet: Learning to Generate Soft-Labels from Noisy-Labels
- Gaussian Universality of Perceptrons with Random Labels
- Dirichlet-Based Prediction Calibration for Learning with Noisy Labels
- Jack and Masters of all Trades: One-Pass Learning Sets of Model Sets From Large Pre-Trained Models
- Variational Rectification Inference for Learning with Noisy Labels
- Riemannian Low-Rank Model Compression for Federated Learning with Over-the-Air Aggregation
- Asymmetric Co-Teaching for Unsupervised Cross Domain Person Re-Identification
- Group privacy for personalized federated learning
- Narrowing the Gap: Improved Detector Training with Noisy Location Annotations
- Unsupervised Clustering and Performance Prediction of Vortex Wakes from Bio-inspired Propulsors
- Visualization Of Class Activation Maps To Explain AI Classification Of Network Packet Captures
- Prediction Model for Mortality Analysis of Pregnant Women Affected With COVID-19
- Identification of diffracted vortex beams at different propagation distances using deep learning
- Weak-shot Fine-grained Classification via Similarity Transfer
- Risk Adversarial Learning System for Connected and Autonomous Vehicle Charging
- FisHook -- An Optimized Approach to Marine Specie Classification using MobileNetV2
- Supervised Contrastive Learning with Nearest Neighbor Search for Speech Emotion Recognition
- Complementary to Multiple Labels: A Correlation-Aware Correction Approach
- Learning the Precise Feature for Cluster Assignment
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Label-Noise Robust Generative Adversarial Networks
- Training Progressively Binarizing Deep Networks Using FPGAs
- Bayesian Statistics Guided Label Refurbishment Mechanism: Mitigating Label Noise in Medical Image Classification
- The Resistance to Label Noise in K-NN and DNN Depends on its Concentration
- BiaSwap: Removing dataset bias with bias-tailored swapping augmentation
- Knodle: Modular Weakly Supervised Learning with PyTorch
- On the Learning Property of Logistic and Softmax Losses for Deep Neural Networks
- Learning Classifiers on Positive and Unlabeled Data with Policy Gradient
- MRCBert: A Machine Reading ComprehensionApproach for Unsupervised Summarization
- Robust Sensible Adversarial Learning of Deep Neural Networks for Image Classification
- ALEX: Active Learning based Enhancement of a Model's Explainability
- A labeled dataset of cloud types using data from GOES-16 and CloudSat
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Ghost Loss to Question the Reliability of Training Data
- Exclusion and Inclusion -- A model agnostic approach to feature importance in DNNs
- UNICS: Multilingual Code Search via Unified Pseudocode and Contrastive Transfer Learning
- CRCEN: A Generalized Cost-sensitive Neural Network Approach for Imbalanced Classification
- Mitigating Memorization in Sample Selection for Learning with Noisy Labels
- The Best of Both Worlds: a Framework for Combining Degradation Prediction with High Performance Super-Resolution Networks
- Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data Fusion
- Mixing between the Cross Entropy and the Expectation Loss Terms
- Paths of A Million People: Extracting Life Trajectories from Wikipedia