Data Programming: Creating Large Training Sets, Quickly
arXiv:1605.07723
Abstract
Large labeled training sets are the critical building blocks of supervised learning methods and are key enablers of deep learning techniques. For some applications, creating labeled training sets is the most time-consuming and expensive part of applying machine learning. We therefore propose a paradigm for the programmatic creation of training sets called data programming in which users express weak supervision strategies or domain heuristics as labeling functions, which are programs that label subsets of the data, but that are noisy and may conflict. We show that by explicitly representing this training set labeling process as a generative model, we can "denoise" the generated training set, and establish theoretically that we can recover the parameters of these generative models in a handful of settings. We then show how to modify a discriminative loss function to make it noise-aware, and demonstrate our method over a range of discriminative models including logistic regression and LSTMs. Experimentally, on the 2014 TAC-KBP Slot Filling challenge, we show that data programming would have led to a new winning score, and also show that applying data programming to an LSTM model leads to a TAC-KBP score almost 6 F1 points over a state-of-the-art LSTM baseline (and into second place in the competition). Additionally, in initial user studies we observed that data programming may be an easier way for non-experts to create machine learning models when training data is limited or unavailable.
References in corpus (1)
Cited by in corpus (48)
- Snorkel: Rapid Training Data Creation with Weak Supervision
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Ontology-driven weak supervision for clinical entity classification in electronic health records
- SwellShark: A Generative Model for Biomedical Named Entity Recognition without Labeled Data
- Fonduer: Knowledge Base Construction from Richly Formatted Data
- StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow
- Drug repurposing for COVID-19 using graph neural network and harmonizing multiple evidence
- Discovering and Validating AI Errors With Crowdsourced Failure Reports
- HoloClean: Holistic Data Repairs with Probabilistic Inference
- OSLNet: Deep Small-Sample Classification with an Orthogonal Softmax Layer
- Generating Multi-Agent Trajectories using Programmatic Weak Supervision
- Deep Convolutional Neural Network Design Patterns
- Denoising Multi-Source Weak Supervision for Neural Text Classification
- Automating Data Science: Prospects and Challenges
- Unsupervised classification to improve the quality of a bird song recording dataset
- Learning from Noisy Labels for Entity-Centric Information Extraction
- Avoiding Your Teacher's Mistakes: Training Neural Networks with Controlled Weak Supervision
- Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
- GOGGLES: Automatic Image Labeling with Affinity Coding
- Reuse and Adaptation for Entity Resolution through Transfer Learning
- Heterogeneous Supervision for Relation Extraction: A Representation Learning Approach
- ORCAS-I: Queries Annotated with Intent using Weak Supervision
- Passage Ranking with Weak Supervision
- Infrastructure for Usable Machine Learning: The Stanford DAWN Project
- Sparse Conditional Hidden Markov Model for Weakly Supervised Named Entity Recognition
- Inspector Gadget: A Data Programming-based Labeling System for Industrial Images
- Link Prediction in Networks with Core-Fringe Data
- Enabling Collaborative Data Science Development with the Ballet Framework
- CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web
- Fidelity-Weighted Learning
- EZLearn: Exploiting Organic Supervision in Large-Scale Data Annotation
- Data Programming using Continuous and Quality-Guided Labeling Functions
- Training Complex Models with Multi-Task Weak Supervision
- Quantitative Overfitting Management for Human-in-the-loop ML Application Development with ease.ml/meter
- Weakly Supervised Label Learning Flows
- Demonstration of Panda: A Weakly Supervised Entity Matching System
- A Deep Representation Empowered Distant Supervision Paradigm for Clinical Information Extraction
- Adversarial Constraint Learning for Structured Prediction
- The Corrective Commit Probability Code Quality Metric
- Follow Your Nose -- Which Code Smells are Worth Chasing?
- Efficient Large-Scale Domain Classification with Personalized Attention
- Machine Learning Systems for Intelligent Services in the IoT: A Survey
- Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence Labeler
- Accurate, Data-Efficient Learning from Noisy, Choice-Based Labels for Inherent Risk Scoring
- Gradual Machine Learning for Entity Resolution
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
- DeepErase: Weakly Supervised Ink Artifact Removal in Document Text Images