Learning with a Wasserstein Loss
arXiv:1506.05439
Abstract
Learning to predict multi-label outputs is challenging, but in many problems there is a natural metric on the outputs that can be used to improve predictions. In this paper we develop a loss function for multi-label learning, based on the Wasserstein distance. The Wasserstein distance provides a natural notion of dissimilarity for probability measures. Although optimizing with respect to the exact Wasserstein distance is costly, recent work has described a regularized approximation that is efficiently computed. We describe an efficient learning algorithm based on this regularization, as well as a novel extension of the Wasserstein distance from probability measures to unnormalized measures. We also describe a statistical learning bound for the loss. The Wasserstein loss can encourage smoothness of the predictions with respect to a chosen metric on the output space. We demonstrate this property on a real-data tag prediction problem, using the Yahoo Flickr Creative Commons dataset, outperforming a baseline that doesn't use the metric.
NIPS 2015; v3 updates Algorithm 1 and Equations 6, 8
References in corpus (2)
Cited by in corpus (110)
- Robust Wasserstein Profile Inference and Applications to Machine Learning
- Stochastic Optimization for Large-scale Optimal Transport
- Soft-DTW: a Differentiable Loss Function for Time-Series
- Estimating individual treatment effect: generalization bounds and algorithms
- Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
- Wasserstein Discriminant Analysis
- Large-Scale Optimal Transport and Mapping Estimation
- Scaling Algorithms for Unbalanced Transport Problems
- Faster Wasserstein Distance Estimation with the Sinkhorn Divergence
- Wasserstein Weisfeiler-Lehman Graph Kernels
- Unnormalized Optimal Transport
- Sliced-Wasserstein Autoencoder: An Embarrassingly Simple Generative Model
- Missing Data Imputation using Optimal Transport
- Sinkhorn Divergences for Unbalanced Optimal Transport
- Entropic Optimal Transport between Unbalanced Gaussian Measures has a Closed Form
- Robust Optimal Transport with Applications in Generative Modeling and Domain Adaptation
- Statistical Inference for Generative Models with Maximum Mean Discrepancy
- Quantum Optimal Transport for Tensor Field Processing
- The Unbalanced Gromov Wasserstein Distance: Conic Formulation and Relaxation
- Rates of Estimation of Optimal Transport Maps using Plug-in Estimators via Barycentric Projections
- Sinkhorn AutoEncoders
- Unbalanced Multi-Marginal Optimal Transport
- On Unbalanced Optimal Transport: An Analysis of Sinkhorn Algorithm
- A mathematical theory of cooperative communication
- Screening Sinkhorn Algorithm for Regularized Optimal Transport
- Coupled VAE: Improved Accuracy and Robustness of a Variational Autoencoder
- Learning with minibatch Wasserstein : asymptotic and gradient properties
- Semi-Supervised Multi-Modal Multi-Instance Multi-Label Deep Network with Optimal Transport
- Minibatch optimal transport distances; analysis and applications
- On the extremal points of the ball of the Benamou-Brenier energy
- Conservative Wasserstein Training for Pose Estimation
- Greedy stochastic algorithms for entropy-regularized optimal transport problems
- Minimum Stein Discrepancy Estimators
- Wasserstein regularization for sparse multi-task regression
- Multilevel Optimal Transport: a Fast Approximation of Wasserstein-1 distances
- Ground Metric Learning on Graphs
- Blind Source Separation with Optimal Transport Non-negative Matrix Factorization
- Scalable Unbalanced Optimal Transport using Generative Adversarial Networks
- Deep multi-class learning from label proportions
- Optimal Feature Transport for Cross-View Image Geo-Localization
- Towards a mathematical theory of trajectory inference
- Gaussian Word Embedding with a Wasserstein Distance Loss
- Improved Deep Spectral Convolution Network For Hyperspectral Unmixing With Multinomial Mixture Kernel and Endmember Uncertainty
- Optimal transport natural gradient for statistical manifolds with continuous sample space
- Information-geometry of physics-informed statistical manifolds and its use in data assimilation
- On the Computation of Kantorovich-Wasserstein Distances between 2D-Histograms by Uncapacitated Minimum Cost Flows
- Evaluating the Disentanglement of Deep Generative Models through Manifold Topology
- Optimal Transport losses and Sinkhorn algorithm with general convex regularization
- Quantum Wasserstein isometries on the qubit state space
- Regularized Optimal Transport is Ground Cost Adversarial
- Regularizing activations in neural networks via distribution matching with the Wasserstein metric
- Wasserstein Style Transfer
- Equivalence Between Wasserstein and Value-Aware Loss for Model-based Reinforcement Learning
- On Scalable and Efficient Computation of Large Scale Optimal Transport
- Reinforced Wasserstein Training for Severity-Aware Semantic Segmentation in Autonomous Driving
- Concentration bounds for linear Monge mapping estimation and optimal transport domain adaptation
- Block-coordinate Frank-Wolfe algorithm and convergence analysis for semi-relaxed optimal transport problem
- A General Framework for Consistent Structured Prediction with Implicit Loss Embeddings
- Statistical Optimal Transport posed as Learning Kernel Embedding
- Entropy-Transport distances between unbalanced metric measure spaces
- Interior-Point Methods Strike Back: Solving the Wasserstein Barycenter Problem
- A contribution to Optimal Transport on incomparable spaces
- The K-Nearest Neighbour UCB algorithm for multi-armed bandits with covariates
- Computing Kantorovich-Wasserstein Distances on -dimensional histograms using -partite graphs
- Regularization Helps with Mitigating Poisoning Attacks: Distributionally-Robust Machine Learning Using the Wasserstein Distance
- Neural Topic Model via Optimal Transport
- Distance Measure Machines
- Unbalanced Optimal Transport through Non-negative Penalized Linear Regression
- Optimal Transport for Stationary Markov Chains via Policy Iteration
- On Linear Optimization over Wasserstein Balls
- Natural gradient via optimal transport
- Neural Network Encapsulation
- Cycle Consistent Probability Divergences Across Different Spaces
- Improving Neural Topic Models with Wasserstein Knowledge Distillation
- Geometric Losses for Distributional Learning
- Regularity as Regularization: Smooth and Strongly Convex Brenier Potentials in Optimal Transport
- Faster Unbalanced Optimal Transport: Translation invariant Sinkhorn and 1-D Frank-Wolfe
- Fast block-coordinate Frank-Wolfe algorithm for semi-relaxed optimal transport
- Online Sinkhorn: Optimal Transport distances from sample streams
- Picture-to-Amount (PITA): Predicting Relative Ingredient Amounts from Food Images
- Utility/Privacy Trade-off through the lens of Optimal Transport
- Matching Guided Distillation
- Scaling positive random matrices: concentration and asymptotic convergence
- Entropy Partial Transport with Tree Metrics: Theory and Practice
- Solving graph compression via optimal transport
- Learning Generalized Gumbel-max Causal Mechanisms
- Manifold optimization for non-linear optimal transport problems
- Matching Distributions via Optimal Transport for Semi-Supervised Learning
- Importance-Aware Semantic Segmentation in Self-Driving with Discrete Wasserstein Training
- Learning to rank for censored survival data
- Learning Where to Look While Tracking Instruments in Robot-assisted Surgery
- Mosaic: A Sample-Based Database System for Open World Query Processing
- CWAE-IRL: Formulating a supervised approach to Inverse Reinforcement Learning problem
- Posterior Ratio Estimation of Latent Variables
- Discovering Invariances in Healthcare Neural Networks
- Coupling Matrix Manifolds and Their Applications in Optimal Transport
- Wasserstein statistics in 1D location-scale model
- Fast Unbalanced Optimal Transport on a Tree
- Generating and Aligning from Data Geometries with Generative Adversarial Networks
- Permutation invariant networks to learn Wasserstein metrics
- Wasserstein Statistics in One-dimensional Location-Scale Model
- Heterogeneous Wasserstein Discrepancy for Incomparable Distributions
- Optimal transport problems regularized by generic convex functions: A geometric and algorithmic approach
- A Functional Perspective on Learning Symmetric Functions with Neural Networks
- Entropy-regularized optimal transport on multivariate normal and q-normal distributions
- Dual Regularized Optimal Transport
- On Multimarginal Partial Optimal Transport: Equivalent Forms and Computational Complexity
- Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk Minimization
- The Fourier Discrepancy Function
- A Linear Transportation Distance for Pattern Recognition