On Missing Labels, Long-tails and Propensities in Extreme Multi-label Classification
arXiv:2207.13186 · doi:10.1145/3534678.3539466
Abstract
The propensity model introduced by Jain et al. 2016 has become a standard approach for dealing with missing and long-tail labels in extreme multi-label classification (XMLC). In this paper, we critically revise this approach showing that despite its theoretical soundness, its application in contemporary XMLC works is debatable. We exhaustively discuss the flaws of the propensity-based approach, and present several recipes, some of them related to solutions used in search engines and recommender systems, that we believe constitute promising alternatives to be followed in XMLC.
This is the author's version of the work accepted at KDD '22
References in corpus (5)
- Conditional Probability Tree Estimation Analysis and Algorithms
- DiSMEC - Distributed Sparse Machines for Extreme Multi-label Classification
- Extreme Classification in Log Memory using Count-Min Sketch: A Case Study of Amazon Search with 50M Products
- Propensity-scored Probabilistic Label Trees
- Unbiased Loss Functions for Multilabel Classification with Missing Labels