23 citations · 28 across the 6 of their papers we have counts for
5 papers · 1 filter
Visual Template Inference for Data Extraction from Documents
Yiming Lin, Mawil Hasan, Rohan Kosalge +2
Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and p…
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Shreya Shankar, Tristan Chambers, Tarak Shah +2
Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for decl…
Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle
Rolando Garcia, Pragya Kallanagoudar, Chithra Anand +4
In this paper we present techniques to incrementally harvest and query arbitrary metadata from machine learning pipelines, without disrupting agile practices. We center our approac…
Towards Accurate and Efficient Document Analytics with Large Language Models
Yiming Lin, Madelon Hulsebos, Ruiying Ma +4
Unstructured data formats account for over 80% of the data currently stored, and extracting value from such formats remains a considerable challenge. In particular, current approac…
SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines
Shreya Shankar, Haotian Li, Parth Asawa +7
Large language models (LLMs) are being increasingly deployed as part of pipelines that repeatedly process or generate data of some sort. However, a common barrier to deployment are…