1 citations · 1 across the 3 of their papers we have counts for
5 papers
Scout: Scalable Document Extraction via Data Similarity
Yiming Lin, Chiyu Hao, Shreya Shankar +1
Extracting values from large document collections powers data analysis across many domains. Frontier LLMs extract such values accurately, but processing an entire collection with o…
Visual Template Inference for Data Extraction from Documents
Yiming Lin, Mawil Hasan, Rohan Kosalge +2
Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and p…
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
Bangrui Xu, Ziyang Miao, Xuanhe Zhou +7
VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together wi…
Can AI Agents Answer Your Data Questions? A Benchmark for Data Agents
Ruiying Ma, Shreya Shankar, Ruiqi Chen +7
Users across enterprises increasingly rely on AI agents to query their data through natural language. However, building reliable data agents remains difficult because real-world da…
LLM-Powered Proactive Data Systems
Sepanta Zeighami, Yiming Lin, Shreya Shankar +1
With the power of LLMs, we now have the ability to query data that was previously impossible to query, including text, images, and video. However, despite this enormous potential,…