Learning the Structure of Generative Models without Labeled Data
arXiv:1703.00854
Abstract
Curating labeled training data has become the primary bottleneck in machine learning. Recent frameworks address this bottleneck with generative models to synthesize labels at scale from weak supervision sources. The generative model's dependency structure directly affects the quality of the estimated labels, but selecting a structure automatically without any labeled data is a distinct challenge. We propose a structure estimation method that maximizes the -regularized marginal pseudolikelihood of the observed data. Our analysis shows that the amount of unlabeled data required to identify the true structure scales sublinearly in the number of possible dependencies for a broad class of models. Simulations show that our method is 100 faster than a maximum likelihood approach and selects as many extraneous dependencies. We also show that our method provides an average of 1.5 F1 points of improvement over existing, user-developed information extraction applications on real-world data such as PubMed journal abstracts.
References in corpus (1)
Cited by in corpus (30)
- Snorkel: Rapid Training Data Creation with Weak Supervision
- A Survey on Data Collection for Machine Learning: a Big Data -- AI Integration Perspective
- Ontology-driven weak supervision for clinical entity classification in electronic health records
- Generating Multi-Agent Trajectories using Programmatic Weak Supervision
- WRENCH: A Comprehensive Benchmark for Weak Supervision
- Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
- LIREx: Augmenting Language Inference with Relevant Explanation
- Deep Probabilistic Logic: A Unifying Framework for Indirect Supervision
- Interactive Weak Supervision: Learning Useful Heuristics for Data Labeling
- Training Complex Models with Multi-Task Weak Supervision
- Knowledge-Based Distant Regularization in Learning Probabilistic Models
- Multi-Resolution Weak Supervision for Sequential Data
- Knodle: Modular Weakly Supervised Learning with PyTorch
- Data Programming by Demonstration: A Framework for Interactively Learning Labeling Functions
- DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation Extraction
- Dependency Structure Misspecification in Multi-Source Weak Supervision Models
- Pairwise Feedback for Data Programming
- FLAME: A Self-Adaptive Auto-labeling System for Heterogeneous Mobile Processors
- Knowledge Efficient Deep Learning for Natural Language Processing
- Combining Probabilistic Logic and Deep Learning for Self-Supervised Learning
- Weakly Supervised Named Entity Tagging with Learnable Logical Rules
- Learning from Multiple Noisy Partial Labelers
- Learning Calibratable Policies using Programmatic Style-Consistency
- Heuristic-Based Weak Learning for Automated Decision-Making
- End-to-End Weak Supervision
- Gradual Machine Learning for Entity Resolution
- Proceedings of the First Workshop on Weakly Supervised Learning (WeaSuL)
- Self-supervised self-supervision by combining deep learning and probabilistic logic
- Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling
- GLaRA: Graph-based Labeling Rule Augmentation for Weakly Supervised Named Entity Recognition