activity
20242026
most citedVisual Template Inference for Data Extraction from Documents

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.DB20261 cited

Visual Template Inference for Data Extraction from Documents

Yiming Lin, Mawil Hasan, Rohan Kosalge +2

Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and p…

cs.HC2025

Steering Semantic Data Processing With DocWrangler

Shreya Shankar, Bhavya Chopra, Mawil Hasan +5

Unstructured text has long been difficult to automatically analyze at scale. Large language models (LLMs) now offer a way forward by enabling {\em semantic data processing}, where…

cs.CL2025

PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines

Reya Vir, Shreya Shankar, Harrison Chase +2

Large language models (LLMs) are increasingly deployed in specialized production data processing pipelines across diverse domains -- such as finance, marketing, and e-commerce. How…

cs.DB2025

DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing

Shreya Shankar, Tristan Chambers, Tarak Shah +2

Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for decl…

cs.DB2024

Flow with FlorDB: Incremental Context Maintenance for the Machine Learning Lifecycle

Rolando Garcia, Pragya Kallanagoudar, Chithra Anand +4

In this paper we present techniques to incrementally harvest and query arbitrary metadata from machine learning pipelines, without disrupting agile practices. We center our approac…