activity
20242026
collaborators
Showing cs.DBShow all

5 papers · 1 filter

cs.DB2026

Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries

Matthew Russo, Yash Agarwal, Tianyu Li +5

Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However, the latter operates as an opa…

cs.DB2026

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing

Xinzhi Wang, Peter Baile Chen, Gerardo Vitagliano +5

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is of…

cs.DB2026

SemBench: A Benchmark for Semantic Query Processing Engines

Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12

We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…

cs.DB2026

KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

Eugenie Lai, Gerardo Vitagliano, Ziyu Zhang +16

Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from ex…

cs.DB20261 cited

Abacus: A Cost-Based Optimizer for Semantic Operator Systems

Matthew Russo, Chunwei Liu, Sivaprasad Sudhir +4

LLMs enable an exciting new class of data processing applications over large collections of unstructured documents. Several new programming frameworks have enabled developers to bu…