Snorkel: Rapid Training Data Creation with Weak Supervision
arXiv:1711.10160 · doi:10.14778/3157794.3157797
Abstract
Labeling training data is increasingly the largest bottleneck in deploying machine learning systems. We present Snorkel, a first-of-its-kind system that enables users to train state-of-the-art models without hand labeling any training data. Instead, users write labeling functions that express arbitrary heuristics, which can have unknown accuracies and correlations. Snorkel denoises their outputs without access to ground truth by incorporating the first end-to-end implementation of our recently proposed machine learning paradigm, data programming. We present a flexible interface layer for writing labeling functions based on our experience over the past year collaborating with companies, agencies, and research labs. In a user study, subject matter experts build models 2.8x faster and increase predictive performance an average 45.5% versus seven hours of hand labeling. We study the modeling tradeoffs in this new setting and propose an optimizer for automating tradeoff decisions that gives up to 1.8x speedup per pipeline execution. In two collaborations, with the U.S. Department of Veterans Affairs and the U.S. Food and Drug Administration, and on four open-source text and image data sets representative of other deployments, Snorkel provides 132% average improvements to predictive performance over prior heuristic approaches and comes within an average 3.60% of the predictive performance of large hand-curated training sets.
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- TensorFlow: A system for large-scale machine learning
- Deep Residual Learning for Image Recognition
- Data Programming: Creating Large Training Sets, Quickly
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Ranking and combining multiple predictors without labeled data
- Learning the Structure of Generative Models without Labeled Data
- Spectral Methods meet EM: A Provably Optimal Algorithm for Crowdsourcing
- Fonduer: Knowledge Base Construction from Richly Formatted Data
- HoloClean: Holistic Data Repairs with Probabilistic Inference
- Inferring Generative Model Structure with Static Analysis
- Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
Cited by in corpus (170)
- On the Opportunities and Risks of Foundation Models
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- A Unifying Review of Deep and Shallow Anomaly Detection
- A Survey on Aspect-Based Sentiment Classification
- A Survey on Data Collection for Machine Learning: a Big Data -- AI Integration Perspective
- A general-purpose material property data extraction pipeline from large polymer corpora using Natural Language Processing
- HoloDetect: Few-Shot Learning for Error Detection
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Ontology-driven weak supervision for clinical entity classification in electronic health records
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Fonduer: Knowledge Base Construction from Richly Formatted Data
- Language Models are Open Knowledge Graphs
- Automatic Detection of Influential Actors in Disinformation Networks
- Scanner: Efficient Video Analysis at Scale
- A Survey of Deep Learning for Scientific Discovery
- Small Sample Learning in Big Data Era
- Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
- Generating Multi-Agent Trajectories using Programmatic Weak Supervision
- Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African Languages
- Denoising Multi-Source Weak Supervision for Neural Text Classification
- Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
- A Data Quality-Driven View of MLOps
- CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT
- WRENCH: A Comprehensive Benchmark for Weak Supervision
- Increasing the Speed and Accuracy of Data LabelingThrough an AI Assisted Interface
- Scim: Intelligent Skimming Support for Scientific Papers
- LIFT: Reinforcement Learning in Computer Systems by Learning From Demonstrations
- End-to-End Entity Resolution for Big Data: A Survey
- Confidence Scores Make Instance-dependent Label-noise Learning Possible
- Ten Steps to Becoming a Musculoskeletal Simulation Expert: A Half-Century of Progress and Outlook for the Future
- OneLabeler: A Flexible System for Building Data Labeling Tools
- E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
- Efficient End-to-End AutoML via Scalable Search Space Decomposition
- Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
- Overton: A Data System for Monitoring and Improving Machine-Learned Products
- ScatterShot: Interactive In-context Example Curation for Text Transformation
- "We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning
- Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
- GOGGLES: Automatic Image Labeling with Affinity Coding
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition
- Rekall: Specifying Video Events using Compositions of Spatiotemporal Labels
- CheXpert++: Approximating the CheXpert labeler for Speed,Differentiability, and Probabilistic Output
- ORCAS-I: Queries Annotated with Intent using Weak Supervision
- Learning to Interpret Satellite Images in Global Scale Using Wikipedia
- VideoPro: A Visual Analytics Approach for Interactive Video Programming
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet Methods
- Fair Generative Modeling via Weak Supervision
- Beyond MeSH: Fine-Grained Semantic Indexing of Biomedical Literature based on Weak Supervision
- Sparse Conditional Hidden Markov Model for Weakly Supervised Named Entity Recognition
- LIREx: Augmenting Language Inference with Relevant Explanation
- Biomedical Named Entity Recognition via Reference-Set Augmented Bootstrapping
- Zero-Shot Learning with Common Sense Knowledge Graphs
- LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
- ODIN: Automated Drift Detection and Recovery in Video Analytics
- Zeus: Efficiently Localizing Actions in Videos using Reinforcement Learning
- In-N-Out: Pre-Training and Self-Training using Auxiliary Information for Out-of-Distribution Robustness
- Toward Code Generation: A Survey and Lessons from Semantic Parsing
- Link Prediction in Networks with Core-Fringe Data
- Tab2Know: Building a Knowledge Base from Tables in Scientific Papers
- Interactive Weak Supervision: Learning Useful Heuristics for Data Labeling
- PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual Data
- Leveraging Multi-Source Weak Social Supervision for Early Detection of Fake News
- SynthBio: A Case Study in Human-AI Collaborative Curation of Text Datasets
- Transformers for End-to-End InfoSec Tasks: A Feasibility Study
- Learning to Interpret Satellite Images Using Wikipedia
- Extracting actionable information from microtexts
- Learning More From Less: Towards Strengthening Weak Supervision for Ad-Hoc Retrieval
- Argument Identification in Public Comments from eRulemaking
- Helix: Holistic Optimization for Accelerating Iterative Machine Learning
- Data Programming using Continuous and Quality-Guided Labeling Functions
- Leveraging Organizational Resources to Adapt Models to New Data Modalities
- Record fusion: A learning approach
- LARCH: Large Language Model-based Automatic Readme Creation with Heuristics
- IoT Virtualization with ML-based Information Extraction
- AutoML to Date and Beyond: Challenges and Opportunities
- Training Complex Models with Multi-Task Weak Supervision
- Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach
- Rapid Image Labeling via Neuro-Symbolic Learning
- Multi-class Text Classification using BERT-based Active Learning
- CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties (Technical Report)
- Prior Adaptive Semi-supervised Learning with Application to EHR Phenotyping
- On Preempting Advanced Persistent Threats Using Probabilistic Graphical Models
- CrossTrainer: Practical Domain Adaptation with Loss Reweighting
- Towards Artificial Intelligence Enabled Financial Crime Detection
- The Complexity of Aggregates over Extractions by Regular Expressions
- End User Authoring of Personalized Content Classifiers: Comparing Example Labeling, Rule Writing, and LLM Prompting
- Quantitative Overfitting Management for Human-in-the-loop ML Application Development with ease.ml/meter
- Data Programming by Demonstration: A Framework for Interactively Learning Labeling Functions
- Train and You'll Miss It: Interactive Model Iteration with Weak Supervision and Pre-Trained Embeddings
- Demonstration of Panda: A Weakly Supervised Entity Matching System
- Knodle: Modular Weakly Supervised Learning with PyTorch
- ANEA: Distant Supervision for Low-Resource Named Entity Recognition
- A Survey of Embedding Space Alignment Methods for Language and Knowledge Graphs
- Semi-Supervised Learning with Declaratively Specified Entropy Constraints
- Learning to Denoise Distantly-Labeled Data for Entity Typing
- Factorized Graph Representations for Semi-Supervised Learning from Sparse Data
- Automated Query-Product Relevance Labeling using Large Language Models for E-commerce Search
- Large-scale investigation of weakly-supervised deep learning for the fine-grained semantic indexing of biomedical literature
- Multi-Resolution Weak Supervision for Sequential Data
- Extracting Semantics from Maintenance Records
- Challenges for cognitive decoding using deep learning methods
- DRo: A data-scarce mechanism to revolutionize the performance of Deep Learning based Security Systems
- Distant Supervision and Noisy Label Learning for Low Resource Named Entity Recognition: A Study on Hausa and Yorùbá
- Learning from Imperfect Annotations
- Data Curation with Deep Learning [Vision]
- Multiple Sclerosis Severity Classification From Clinical Text
- Importance Reweighting for Biquality Learning
- Global Multiclass Classification and Dataset Construction via Heterogeneous Local Experts
- Technical Report on Data Integration and Preparation
- Dependency Structure Misspecification in Multi-Source Weak Supervision Models
- DIAG-NRE: A Neural Pattern Diagnosis Framework for Distantly Supervised Neural Relation Extraction
- Named Entity Recognition with Partially Annotated Training Data
- Clarify: Improving Model Robustness With Natural Language Corrections
- Creating Training Sets via Weak Indirect Supervision
- Data augmentation on graphs for table type classification
- AutoNLU: Detecting, root-causing, and fixing NLU model errors
- Weight Annotation in Information Extraction
- Automated Label Generation for Time Series Classification with Representation Learning: Reduction of Label Cost for Training
- Textbook to triples: Creating knowledge graph in the form of triples from AI TextBook
- Training Transformers for Information Security Tasks: A Case Study on Malicious URL Prediction
- Whither AutoML? Understanding the Role of Automation in Machine Learning Workflows
- Summarizing Text on Any Aspects: A Knowledge-Informed Weakly-Supervised Approach
- Adaptive Rule Discovery for Labeling Text Data
- Detecting Fake News with Weak Social Supervision
- MobIE: A German Dataset for Named Entity Recognition, Entity Linking and Relation Extraction in the Mobility Domain
- Learning from Multiple Noisy Partial Labelers
- OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning
- BERTifying the Hidden Markov Model for Multi-Source Weakly Supervised Named Entity Recognition
- Online publication of court records: circumventing the privacy-transparency trade-off
- Towards Automated Evaluation of Explanations in Graph Neural Networks
- Dynamic Mode Decomposition based feature for Image Classification
- Theoretical Model and Practical Considerations for Data Lineage Reconstruction
- FLAME: A Self-Adaptive Auto-labeling System for Heterogeneous Mobile Processors
- Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
- Weakly Supervised Named Entity Tagging with Learnable Logical Rules
- Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence Labeler
- Who is we? Disambiguating the referents of first person plural pronouns in parliamentary debates
- Search4Code: Code Search Intent Classification Using Weak Supervision
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models
- TAGLETS: A System for Automatic Semi-Supervised Learning with Auxiliary Data
- Modular Self-Supervision for Document-Level Relation Extraction
- LinTO : Assistant vocal open-source respectueux des données personnelles pour les réunions d'entreprise
- Named Entity Recognition -- Is there a glass ceiling?
- SentiQ: A Probabilistic Logic Approach to Enhance Sentiment Analysis Tool Quality
- Training Machine Learning Models by Regularizing their Explanations
- Noisy Labels for Weakly Supervised Gamma Hadron Classification
- DomiKnowS: A Library for Integration of Symbolic Domain Knowledge in Deep Learning
- HERALD: An Annotation Efficient Method to Detect User Disengagement in Social Conversations
- Representation Learning from Limited Educational Data with Crowdsourced Labels
- Data Management for Causal Algorithmic Fairness
- TAACKIT: Track Annotation and Analytics with Continuous Knowledge Integration Tool
- Zero-shot Task Transfer for Invoice Extraction via Class-aware QA Ensemble
- Proceedings of the First Workshop on Weakly Supervised Learning (WeaSuL)
- Teach me how to Label: Labeling Functions from Natural Language with Text-to-text Transformers
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
- Fast Online "Next Best Offers" using Deep Learning
- Weak-supervision for Deep Representation Learning under Class Imbalance
- Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling
- Active WeaSuL: Improving Weak Supervision with Active Learning
- Adapting Coreference Resolution for Processing Violent Death Narratives
- Addressing Training Bias via Automated Image Annotation
- Split-Correctness in Information Extraction
- Self-Training with Weak Supervision
- GLaRA: Graph-based Labeling Rule Augmentation for Weakly Supervised Named Entity Recognition
- End-to-End Weak Supervision
- Heuristic-Based Weak Learning for Automated Decision-Making
- Learning Structured Representations of Entity Names using Active Learning and Weak Supervision
- KnowMAN: Weakly Supervised Multinomial Adversarial Networks
- Gradual Machine Learning for Entity Resolution
- Teaching Autoregressive Language Models Complex Tasks By Demonstration