Wasserstein Discriminant Analysis
arXiv:1608.08063 · doi:10.1007/s10994-018-5717-1
Abstract
Wasserstein Discriminant Analysis (WDA) is a new supervised method that can improve classification of high-dimensional data by computing a suitable linear map onto a lower dimensional subspace. Following the blueprint of classical Linear Discriminant Analysis (LDA), WDA selects the projection matrix that maximizes the ratio of two quantities: the dispersion of projected points coming from different classes, divided by the dispersion of projected points coming from the same class. To quantify dispersion, WDA uses regularized Wasserstein distances, rather than cross-variance measures which have been usually considered, notably in LDA. Thanks to the the underlying principles of optimal transport, WDA is able to capture both global (at distribution scale) and local (at samples scale) interactions between classes. Regularized Wasserstein distances can be computed using the Sinkhorn matrix scaling algorithm; We show that the optimization of WDA can be tackled using automatic differentiation of Sinkhorn iterations. Numerical experiments show promising results both in terms of prediction and visualization on toy examples and real life datasets such as MNIST and on deep features obtained from a subset of the Caltech dataset.
References in corpus (1)
Cited by in corpus (18)
- Near-linear time approximation algorithms for optimal transport via Sinkhorn iteration
- Differential Properties of Sinkhorn Approximation for Learning with Wasserstein Distance
- Unsupervised Alignment of Embeddings with Wasserstein Procrustes
- Differentiation and regularity of semi-discrete optimal transport with respect to the parameters of the discrete measure
- Screening Sinkhorn Algorithm for Regularized Optimal Transport
- Quantum Optimal Transport
- Wasserstein Distributionally Robust Optimization: Theory and Applications in Machine Learning
- When OT meets MoM: Robust estimation of Wasserstein Distance
- Quantitative stability of optimal transport maps and linearization of the 2-Wasserstein space
- Deep multi-class learning from label proportions
- Differentiable Ranks and Sorting using Optimal Transport
- Differentiable Particle Filtering via Entropy-Regularized Optimal Transport
- Regularized Optimal Transport is Ground Cost Adversarial
- Stochastic Optimization for Regularized Wasserstein Estimators
- A Unified Joint Maximum Mean Discrepancy for Domain Adaptation
- Tensor optimal transport, distance between sets of measures and tensor scaling
- Metric Learning via Maximizing the Lipschitz Margin Ratio
- Wasserstein t-SNE