Learning from Multiple Noisy Partial Labelers
arXiv:2106.04530
Abstract
Programmatic weak supervision creates models without hand-labeled training data by combining the outputs of heuristic labelers. Existing frameworks make the restrictive assumption that labelers output a single class label. Enabling users to create partial labelers that output subsets of possible class labels would greatly expand the expressivity of programmatic weak supervision. We introduce this capability by defining a probabilistic generative model that can estimate the underlying accuracies of multiple noisy partial labelers without ground truth labels. We show how to scale up learning, for example learning on 100k examples in one minute, a 300x speed up compared to a naive implementation. We also prove that this class of models is generically identifiable up to label swapping under mild conditions. We evaluate our framework on three text classification and six object classification tasks. On text tasks, adding partial labels increases average accuracy by 8.6 percentage points. On image tasks, we show that partial labels allow us to approach some zero-shot object classification problems with programmatic weak supervision by using class attributes as partial labelers. On these tasks, our framework has accuracy comparable to recent embedding-based zero-shot learning methods, while using only pre-trained attribute detectors.
In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics (AISTATS) 2022
References in corpus (16)
- Decoupled Weight Decay Regularization
- Zero-Shot Learning Through Cross-Modal Transfer
- Snorkel: Rapid Training Data Creation with Weak Supervision
- Zero-Shot Learning by Convex Combination of Semantic Embeddings
- Identifiability of parameters in latent structure models with many observed variables
- Zero Shot Recognition with Unreliable Attributes
- Learning the Structure of Generative Models without Labeled Data
- Ontology-driven weak supervision for clinical entity classification in electronic health records
- Learning from Complementary Labels
- Aequitas: A Bias and Fairness Audit Toolkit
- Minimax Optimal Convergence Rates for Estimating Ground Truth from Crowdsourced Labels
- Learning from Rules Generalizing Labeled Exemplars
- Learning Dependency Structures for Weak Supervision Models
- Scalable Semi-Supervised Aggregation of Classifiers
- Attribute Propagation Network for Graph Zero-shot Learning
- Disambiguation of weak supervision with exponential convergence rates