works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
most citedAbacus: A Cost-Based Optimizer for Semantic Operator Systems

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.AI2026

AutoSynthesis: An agentic system for automated meta-analysis

Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano +2

AutoSynthesis is a multi‑agent AI system that takes a natural‑language research question and automatically conducts a full quantitative meta‑analysis, from literature search to eff…

cs.DB2026

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing

Xinzhi Wang, Peter Baile Chen, Gerardo Vitagliano +5

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is of…

cs.DB2026

SemBench: A Benchmark for Semantic Query Processing Engines

Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12

We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…

cs.DB2026

KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

Eugenie Lai, Gerardo Vitagliano, Ziyu Zhang +16

Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from ex…

cs.DB20261 cited

Abacus: A Cost-Based Optimizer for Semantic Operator Systems

Matthew Russo, Chunwei Liu, Sivaprasad Sudhir +4

LLMs enable an exciting new class of data processing applications over large collections of unstructured documents. Several new programming frameworks have enabled developers to bu…

cs.AI2025

PalimpChat: Declarative and Interactive AI analytics

Chunwei Liu, Gerardo Vitagliano, Brandon Rose +3

Thanks to the advances in generative architectures and large language models, data scientists can now code pipelines of machine-learning operations to process large collections of…