1 citations · 1 across the 6 of their papers we have counts for
7 papers · 1 filter
SemJoin: Semantic Join Optimization
Christopher Gou, Aditya Banerjee, Jiaxuan Wang +1
Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis. A semantic join, joining two ta…
TabClean: Reusable LLM-Synthesized Programs for Tabular Data Cleaning
Yibo Wang, Riteng Zhang, Yinghao He +3
Reliable analytics and machine-learning pipelines depend on clean tabular data, yet production tables often contain missing values, typographical errors, inconsistent formats, viol…
SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing
Xinzhi Wang, Peter Baile Chen, Gerardo Vitagliano +5
Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is of…
iPDB -- Optimizing Semantic SQL Queries
Udesh Kumarasinghe, Tyler Liu, Ahmed R. Mahmood +2
Structured Query Language (SQL) has remained the standard query language for databases. SQL is highly optimized for processing structured data laid out in relations. Meanwhile, in…
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
Yan Zhou, Chunwei Liu, Bhuvan Urgaonkar +11
Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query…
Abacus: A Cost-Based Optimizer for Semantic Operator Systems
Matthew Russo, Chunwei Liu, Sivaprasad Sudhir +4
LLMs enable an exciting new class of data processing applications over large collections of unstructured documents. Several new programming frameworks have enabled developers to bu…