786 citations · 861 across the 5 of their papers we have counts for
8 papers
Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Laurel Orr, Megan Leszczynski, Simran Arora +4
A challenge for named entity disambiguation (NED), the task of mapping textual mentions to entities in a knowledge base, is how to disambiguate entities that appear rarely in the t…
Train and You'll Miss It: Interactive Model Iteration with Weak Supervision and Pre-Trained Embeddings
Mayee F. Chen, Daniel Y. Fu, Frederic Sala +5
Our goal is to enable machine learning systems to be trained interactively. This requires models that perform well and train quickly, without large amounts of hand-labeled data. We…
Understanding and Improving Information Transfer in Multi-Task Learning
Sen Wu, Hongyang R. Zhang, Christopher Ré
We investigate multi-task learning approaches that use a shared feature representation for all tasks. To better understand the transfer of task information, we study an architectur…
Ivy: Instrumental Variable Synthesis for Causal Inference
Zhaobin Kuang, Frederic Sala, Nimit Sohoni +5
A popular way to estimate the causal effect of a variable x on y from observational data is to use an instrumental variable (IV): a third variable z that affects y only through x.…
Understanding the Downstream Instability of Word Embeddings
Megan Leszczynski, Avner May, Jian Zhang +3
Many industrial machine learning (ML) systems require frequent retraining to keep up-to-date with constantly changing data. This retraining exacerbates a large challenge facing ML…
Slice-based Learning: A Programming Model for Residual Learning in Critical Data Slices
Vincent S. Chen, Sen Wu, Zhenzhen Weng +2
In real-world machine learning applications, data subsets correspond to especially critical outcomes: vulnerable cyclist detections are safety-critical in an autonomous driving tas…