activity
20242026
most citedPalimpChat: Declarative and Interactive AI analytics

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.DBShow all

7 papers · 1 filter

cs.DB2026

SemJoin: Semantic Join Optimization

Christopher Gou, Aditya Banerjee, Jiaxuan Wang +1

Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis. A semantic join, joining two ta…

cs.DB2026

TabClean: Reusable LLM-Synthesized Programs for Tabular Data Cleaning

Yibo Wang, Riteng Zhang, Yinghao He +3

Reliable analytics and machine-learning pipelines depend on clean tabular data, yet production tables often contain missing values, typographical errors, inconsistent formats, viol…

cs.DB2026

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing

Xinzhi Wang, Peter Baile Chen, Gerardo Vitagliano +5

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is of…

cs.DB2026

iPDB -- Optimizing Semantic SQL Queries

Udesh Kumarasinghe, Tyler Liu, Ahmed R. Mahmood +2

Structured Query Language (SQL) has remained the standard query language for databases. SQL is highly optimized for processing structured data laid out in relations. Meanwhile, in…

cs.DB2025

PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking

Yan Zhou, Chunwei Liu, Bhuvan Urgaonkar +11

Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query…

cs.DB2025

Abacus: A Cost-Based Optimizer for Semantic Operator Systems

Matthew Russo, Chunwei Liu, Sivaprasad Sudhir +4

LLMs enable an exciting new class of data processing applications over large collections of unstructured documents. Several new programming frameworks have enabled developers to bu…