collaborators

9 papers

cs.DB2026

100x Cost & Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models

Yeounoh Chung, Rushabh Desai, Jian He +9

Several data warehouse and database providers have recently introduced extensions to SQL called AI Queries, enabling users to specify functions and conditions in SQL that are evalu…

cs.DB2026

Multi-Objective Agentic Rewrites for Unstructured Data Processing

Lindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami +3

One year ago, we open-sourced DocETL, a declarative system for LLM-powered data processing that, as of March 2026, has 3.7K GitHub stars and users across domains (e.g., journalism,…

cs.DB2026

An In-Depth Study of Filter-Agnostic Vector Search on a PostgreSQL Database System: [Experiments and Analysis]

Duo Lu, Helena Caminal, Manos Chatzakis +4

Filtered Vector Search (FVS) is critical for supporting semantic search and GenAI applications in modern database systems. However, existing research most often evaluates algorithm…

cs.DB2026

SemBench: A Benchmark for Semantic Query Processing Engines

Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12

We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…

cs.IR2026

Fine-Grained Table Retrieval Through the Lens of Complex Queries

Wojciech Kosiuk, Xingyu Ji, Yeounoh Chung +2

Enabling question answering over tables and databases in natural language has become a key capability in the democratization of insights from tabular data sources. These systems fi…

cs.DB2026

High-Fidelity And Complex Test Data Generation For Google SQL Code Generation Services

Shivasankari Kannan, Yeounoh Chung, Amita Gondi +2

The demand for high-fidelity test data is paramount in industrial settings where access to production data is largely restricted. Traditional data generation methods often fall sho…