Large-scale Multi-label Learning with Missing Labels
arXiv:1307.5101
Abstract
The multi-label classification problem has generated significant interest in recent years. However, existing approaches do not adequately address two key challenges: (a) the ability to tackle problems with a large number (say millions) of labels, and (b) the ability to handle data with missing labels. In this paper, we directly address both these problems by studying the multi-label problem in a generic empirical risk minimization (ERM) framework. Our framework, despite being simple, is surprisingly able to encompass several recent label-compression based methods which can be derived as special cases of our method. To optimize the ERM problem, we develop techniques that exploit the structure of specific loss functions - such as the squared loss function - to offer efficient algorithms. We further show that our learning framework admits formal excess risk bounds even in the presence of missing labels. Our risk bounds are tight and demonstrate better generalization performance for low-rank promoting trace-norm regularization when compared to (rank insensitive) Frobenius norm regularization. Finally, we present extensive empirical results on a variety of benchmark datasets and show that our methods perform significantly better than existing label compression based methods and can scale up to very large datasets such as the Wikipedia dataset.
References in corpus (1)
Cited by in corpus (86)
- YouTube-8M: A Large-Scale Video Classification Benchmark
- The Emerging Trends of Multi-Label Learning
- Global Optimality of Local Search for Low Rank Matrix Recovery
- Joint Ranking SVM and Binary Relevance with Robust Low-Rank Learning for Multi-Label Classification
- Bonsai -- Diverse and Shallow Trees for Extreme Multi-label Classification
- PU Learning for Matrix Completion
- Expand Globally, Shrink Locally: Discriminant Multi-label Learning with Missing Labels
- Materials Representation and Transfer Learning for Multi-Property Prediction
- Learning Deep Latent Spaces for Multi-Label Classification
- Dropping Convexity for Faster Semi-definite Optimization
- Logarithmic Time Online Multiclass prediction
- Disentangled Variational Autoencoder based Multi-Label Classification with Covariance-Aware Multivariate Probit Model
- DiSMEC - Distributed Sparse Machines for Extreme Multi-label Classification
- Towards Scalable and Reliable Capsule Networks for Challenging NLP Applications
- PECOS: Prediction for Enormous and Correlated Output Spaces
- Learning a Deep ConvNet for Multi-label Classification with Partial Labels
- A Greedy Approach for Budgeted Maximum Inner Product Search
- Structured Prediction Energy Networks
- Taming Pretrained Transformers for Extreme Multi-label Text Classification
- Towards Label Imbalance in Multi-label Classification with Many Labels
- Streaming Label Learning for Modeling Labels on the Fly
- Zoo Guide to Network Embedding
- Tagging like Humans: Diverse and Distinct Image Annotation
- sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification
- LdSM: Logarithm-depth Streaming Multi-label Decision Trees
- DocTag2Vec: An Embedding Based Multi-label Learning Approach for Document Tagging
- Locally Non-linear Embeddings for Extreme Multi-label Learning
- Marginal loss and exclusion loss for partially supervised multi-organ segmentation
- Deep Extreme Multi-label Learning
- Efficient Loss-Based Decoding on Graphs For Extreme Classification
- Multi-label Learning with Missing Labels using Mixed Dependency Graphs
- Privileged Multi-label Learning
- Label Disentanglement in Partition-based Extreme Multilabel Classification
- Target-Embedding Autoencoders for Supervised Representation Learning
- Unbiased Loss Functions for Extreme Classification With Missing Labels
- Prototypical Networks for Multi-Label Learning
- Subset Labeled LDA for Large-Scale Multi-Label Classification
- SPL-MLL: Selecting Predictable Landmarks for Multi-Label Learning
- Extreme Multi-label Classification from Aggregated Labels
- Learning-to-Rank with Partitioned Preference: Fast Estimation for the Plackett-Luce Model
- On Learning High Dimensional Structured Single Index Models
- HERA: Partial Label Learning by Combining Heterogeneous Loss with Sparse and Low-Rank Regularization
- Nonconvex One-bit Single-label Multi-label Learning
- Doubly-stochastic mining for heterogeneous retrieval
- Sparse Group Inductive Matrix Completion
- Extreme Classification in Log Memory
- Local Rademacher Complexity for Multi-label Learning
- On the benefits of output sparsity for multi-label classification
- RIPML: A Restricted Isometry Property based Approach to Multilabel Learning
- Self-Paced Multi-Label Learning with Diversity
- Knowledge-Based Construction of Confusion Matrices for Multi-Label Classification Algorithms using Semantic Similarity Measures
- Efficient Optimization Methods for Extreme Similarity Learning with Nonlinear Embeddings
- Multilabel Classification by Hierarchical Partitioning and Data-dependent Grouping
- Structured Prediction with Partial Labelling through the Infimum Loss
- Collaborative Graph Walk for Semi-supervised Multi-Label Node Classification
- SepNE: Bringing Separability to Network Embedding
- Aggressive Sampling for Multi-class to Binary Reduction with Applications to Text Classification
- Regret Bounds for Non-decomposable Metrics with Missing Labels
- Block-wise Partitioning for Extreme Multi-label Classification
- Classification of sparse binary vectors
- Ranking-Based Autoencoder for Extreme Multi-label Classification
- Representation learning of drug and disease terms for drug repositioning
- Collaborative Filtering and Multi-Label Classification with Matrix Factorization
- Learning of Generalized Low-Rank Models: A Greedy Approach
- An Efficient Large-scale Semi-supervised Multi-label Classifier Capable of Handling Missing labels
- CCMN: A General Framework for Learning with Class-Conditional Multi-Label Noise
- Overcoming the curse of dimensionality with Laplacian regularization in semi-supervised learning
- On Riemannian Approach for Constrained Optimization Model in Extreme Classification Problems
- Multi-Label Learning with Global and Local Label Correlation
- DeepHelp: Deep Learning for Shout Crisis Text Conversations
- Improving Multi-label Learning with Missing Labels by Structured Semantic Correlations
- Multi-Label Learning with Provable Guarantee
- Group Preserving Label Embedding for Multi-Label Classification
- Deep Hiearchical Multi-Label Classification Applied to Chest X-Ray Abnormality Taxonomies
- Pseudo Labeling and Negative Feedback Learning for Large-scale Multi-label Domain Classification
- Subspace Clustering Based Tag Sharing for Inductive Tag Matrix Refinement with Complex Errors
- On-the-fly Global Embeddings Using Random Projections for Extreme Multi-label Classification
- Data-dependent Generalization Bounds for Multi-class Classification
- A method of supervised learning from conflicting data with hidden contexts
- Learning with Holographic Reduced Representations
- Leveraging Distributional Semantics for Multi-Label Learning
- On Learning Vector Representations in Hierarchical Label Spaces
- Distribution-based Label Space Transformation for Multi-label Learning
- LLC: Accurate, Multi-purpose Learnt Low-dimensional Binary Codes
- Fast Multi-label Learning
- Active Refinement for Multi-Label Learning: A Pseudo-Label Approach