Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
arXiv:1611.05916
Abstract
In the context of single-label classification, despite the huge success of deep learning, the commonly used cross-entropy loss function ignores the intricate inter-class relationships that often exist in real-life tasks such as age classification. In this work, we propose to leverage these relationships between classes by training deep nets with the exact squared Earth Mover's Distance (also known as Wasserstein distance) for single-label classification. The squared EMD loss uses the predicted probabilities of all classes and penalizes the miss-predictions according to a ground distance matrix that quantifies the dissimilarities between classes. We demonstrate that on datasets with strong inter-class relationships such as an ordering between classes, our exact squared EMD losses yield new state-of-the-art results. Furthermore, we propose a method to automatically learn this matrix using the CNN's own features during training. We show that our method can learn a ground distance matrix efficiently with no inter-class relationship priors and yield the same performance gain. Finally, we show that our method can be generalized to applications that lack strong inter-class relationships and still maintain state-of-the-art performance. Therefore, with limited computational overhead, one can always deploy the proposed loss function on any dataset over the conventional cross-entropy.
References in corpus (4)
Cited by in corpus (22)
- Convolutional Ordinal Regression Forest for Image Ordinal Estimation
- Making Images Real Again: A Comprehensive Survey on Deep Image Composition
- Measuring a hate speech spectrum with faceted Rasch item response theory and perspective-aware, explainable-by-design deep learning
- Dim but not entirely dark: Extracting the Galactic Center Excess' source-count distribution with neural nets
- Fine-Grained Age Estimation in the wild with Attention LSTM Networks
- DeepHist: Differentiable Joint and Color Histogram Layers for Image-to-Image Translation
- Leveraging Class Hierarchies with Metric-Guided Prototype Learning
- Unimodal probability distributions for deep ordinal classification
- Age Group and Gender Estimation in the Wild with Deep RoR Architecture
- Theoretical Insights Into Multiclass Classification: A High-dimensional Asymptotic View
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Differentiable Earth Mover's Distance for Data Compression at the High-Luminosity LHC
- Image Composition Assessment with Saliency-augmented Multi-pattern Pooling
- Predicting Aesthetic Score Distribution through Cumulative Jensen-Shannon Divergence
- MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment
- The Earth Mover's Pinball Loss: Quantiles for Histogram-Valued Regression
- Image Aesthetics Prediction Using Multiple Patches Preserving the Original Aspect Ratio of Contents
- Deep Ordinal Regression using Optimal Transport Loss and Unimodal Output Probabilities
- Deep Ordinal Regression for Pledge Specificity Prediction
- Learning to rank for censored survival data
- United We Learn Better: Harvesting Learning Improvements From Class Hierarchies Across Tasks
- Hue-Net: Intensity-based Image-to-Image Translation with Differentiable Histogram Loss Functions