DeepXML: A Deep Extreme Multi-Label Learning Framework Applied to Short Text Documents
arXiv:2111.06685 · doi:10.1145/3437963.3441810
Abstract
Scalability and accuracy are well recognized challenges in deep extreme multi-label learning where the objective is to train architectures for automatically annotating a data point with the most relevant subset of labels from an extremely large label set. This paper develops the DeepXML framework that addresses these challenges by decomposing the deep extreme multi-label task into four simpler sub-tasks each of which can be trained accurately and efficiently. Choosing different components for the four sub-tasks allows DeepXML to generate a family of algorithms with varying trade-offs between accuracy and scalability. In particular, DeepXML yields the Astec algorithm that could be 2-12% more accurate and 5-30x faster to train than leading deep extreme classifiers on publically available short text datasets. Astec could also efficiently train on Bing short text datasets containing up to 62 million labels while making predictions for billions of users and data points per day on commodity hardware. This allowed Astec to be deployed on the Bing search engine for a number of short text applications ranging from matching user queries to advertiser bid phrases to showing personalized ads where it yielded significant gains in click-through-rates, coverage, revenue and other online metrics over state-of-the-art techniques currently in production. DeepXML's code is available at https://github.com/Extreme-classification/deepxml
References in corpus (10)
- Pre-training Tasks for Embedding-based Large-scale Retrieval
- Simrank++: Query rewriting through link analysis of the click graph
- DECAF: Deep Extreme Classification with Label Features
- DiSMEC - Distributed Sparse Machines for Extreme Multi-label Classification
- ECLARE: Extreme Classification with Label Graph Correlations
- Extreme Multi-Label Legal Text Classification: A case study in EU Legislation
- Extreme Classification in Log Memory using Count-Min Sketch: A Case Study of Amazon Search with 50M Products
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text Classification
- Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation
- An end-to-end Generative Retrieval Method for Sponsored Search Engine --Decoding Efficiently into a Closed Target Domain
Cited by in corpus (12)
- DECAF: Deep Extreme Classification with Label Features
- ECLARE: Extreme Classification with Label Graph Correlations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text Classification
- FreshDiskANN: A Fast and Accurate Graph-Based ANN Index for Streaming Similarity Search
- Multi-modal Extreme Classification
- Extreme Meta-Classification for Large-Scale Zero-Shot Retrieval
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- Efficient Text Encoders for Labor Market Analysis
- GraphEx: A Graph-based Extraction Method for Advertiser Keyphrase Recommendation
- BroadGen: A Framework for Generating Effective and Efficient Advertiser Broad Match Keyphrase Recommendations
- Learning with Holographic Reduced Representations
- LLC: Accurate, Multi-purpose Learnt Low-dimensional Binary Codes