75 citations · 78 across the 6 of their papers we have counts for
6 papers
Cross-Domain Data Integration for Named Entity Disambiguation in Biomedical Text
Maya Varma, Laurel Orr, Sen Wu +3
Named entity disambiguation (NED), which involves mapping textual mentions to structured entities, is particularly challenging in the medical domain due to the presence of rare ent…
Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems
Laurel Orr, Atindriyo Sanyal, Xiao Ling +2
The industrial machine learning pipeline requires iterating on model features, training and deploying models, and monitoring deployed models at scale. Feature stores were developed…
Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Laurel Orr, Megan Leszczynski, Simran Arora +4
A challenge for named entity disambiguation (NED), the task of mapping textual mentions to entities in a knowledge base, is how to disambiguate entities that appear rarely in the t…
Sample Debiasing in the Themis Open World Database System (Extended Version)
Laurel Orr, Magda Balazinska, Dan Suciu
Open world database management systems assume tuples not in the database still exist and are becoming an increasingly important area of research. We present Themis, the first open…
Mosaic: A Sample-Based Database System for Open World Query Processing
Laurel Orr, Samuel Ainsworth, Walter Cai +3
Data scientists have relied on samples to analyze populations of interest for decades. Recently, with the increase in the number of public data repositories, sample data has become…
EntropyDB: A Probabilistic Approach to Approximate Query Processing
Laurel Orr, Magdalena Balazinska, Dan Suciu
We present EntropyDB, an interactive data exploration system that uses a probabilistic approach to generate a small, query-able summary of a dataset. Departing from traditional sum…