most cited"We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine Learning

23 citations · 28 across the 6 of their papers we have counts for

collaborators
Showing cs.DBShow all

5 papers · 1 filter

cs.DB20251 cited

Visual Template Inference for Data Extraction from Documents

Yiming Lin, Mawil Hasan, Rohan Kosalge +2

Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and p…

cs.DB20244 cited

DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

Shreya Shankar, Tristan Chambers, Tarak Shah +2

Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for decl…

cs.DB2024

Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle

Rolando Garcia, Pragya Kallanagoudar, Chithra Anand +4

In this paper we present techniques to incrementally harvest and query arbitrary metadata from machine learning pipelines, without disrupting agile practices. We center our approac…

cs.DB20243 cited

Towards Accurate and Efficient Document Analytics with Large Language Models

Yiming Lin, Madelon Hulsebos, Ruiying Ma +4

Unstructured data formats account for over 80% of the data currently stored, and extracting value from such formats remains a considerable challenge. In particular, current approac…

cs.DB2024

SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines

Shreya Shankar, Haotian Li, Parth Asawa +7

Large language models (LLMs) are being increasingly deployed as part of pipelines that repeatedly process or generate data of some sort. However, a common barrier to deployment are…