Training Deep Neural Networks via Direct Loss Minimization
arXiv:1511.06411
Abstract
Supervised training of deep neural nets typically relies on minimizing cross-entropy. However, in many domains, we are interested in performing well on metrics specific to the application. In this paper we propose a direct loss minimization approach to train deep neural networks, which provably minimizes the application-specific loss function. This is often non-trivial, since these functions are neither smooth nor decomposable and thus are not amenable to optimization with standard gradient-based methods. We demonstrate the effectiveness of our approach in the context of maximizing average precision for ranking problems. Towards this goal, we develop a novel dynamic programming algorithm that can efficiently compute the weight updates. Our approach proves superior to a variety of baselines in the context of action classification and object detection, especially in the presence of label noise.
ICML2016
Cited by in corpus (27)
- Towards Accurate One-Stage Object Detection with AP-Loss
- AP-Loss for Accurate One-Stage Object Detection
- Few-Shot Learning Through an Information Retrieval Lens
- Transfer learning for radio galaxy classification
- Optimizing Non-decomposable Measures with Deep Networks
- Learning Surrogate Losses
- End-to-end training of object class detectors for mean average precision
- Auto Seg-Loss: Searching Metric Surrogates for Semantic Segmentation
- Stochastic Optimization of Areas Under Precision-Recall Curves with Provable Convergence
- Implicit MLE: Backpropagating Through Discrete Exponential Family Distributions
- Direct Optimization through for Discrete Variational Auto-Encoder
- Learning Randomly Perturbed Structured Predictors for Direct Loss Minimization
- AutoLoss-Zero: Searching Loss Functions from Scratch for Generic Tasks
- Momentum Accelerates the Convergence of Stochastic AUPRC Maximization
- A review on ranking problems in statistical learning
- AP-Perf: Incorporating Generic Performance Metrics in Differentiable Learning
- VIABLE: Fast Adaptation via Backpropagating Learned Loss
- Direct Policy Gradients: Direct Optimization of Policies in Discrete Action Spaces
- Learning Discriminators as Energy Networks in Adversarial Learning
- Hashing as Tie-Aware Learning to Rank
- SynSig2Vec: Learning Representations from Synthetic Dynamic Signatures for Real-world Verification
- MetricOpt: Learning to Optimize Black-Box Evaluation Metrics
- A Surrogate Objective Framework for Prediction+Optimization with Soft Constraints
- Stateless actor-critic for instance segmentation with high-level priors
- Training Over-parameterized Models with Non-decomposable Objectives
- The Contextual Appointment Scheduling Problem
- A Unified Framework of Surrogate Loss by Refactoring and Interpolation