activity
20142024
most citedData Ambiguity Strikes Back: How Documentation Improves GPT's Text-to-SQL

3 citations · 8 across the 8 of their papers we have counts for

collaborators
Showing cs.DBShow all

7 papers · 1 filter

cs.DB2024

Cocoon: Semantic Table Profiling Using Large Language Models

Zezhou Huang, Eugene Wu

Data profilers play a crucial role in the preprocessing phase of data analysis by identifying quality issues such as missing, extreme, or erroneous values. Traditionally, profilers…

cs.DB2024

SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines

Shreya Shankar, Haotian Li, Parth Asawa +7

Large language models (LLMs) are being increasingly deployed as part of pipelines that repeatedly process or generate data of some sort. However, a common barrier to deployment are…

cs.DB20233 cited

Data Ambiguity Strikes Back: How Documentation Improves GPT's Text-to-SQL

Zezhou Huang, Pavan Kalyan Damalapati, Eugene Wu

Text-to-SQL allows experts to use databases without in-depth knowledge of them. However, real-world tasks have both query and data ambiguities. Most works on Text-to-SQL focused on…

cs.DB2023

Lightweight Materialization for Fast Dashboards Over Joins

Zezhou Huang, Eugene Wu

Dashboards are vital in modern business intelligence tools, providing non-technical users with an interface to access comprehensive business data. With the rise of cloud technology…

cs.DB20232 cited

The Fast and the Private: Task-based Dataset Search

Zezhou Huang, Jiaxiang Liu, Haonan Wang +1

Modern dataset search platforms employ ML task-based utility metrics instead of relying on metadata-based keywords to comb through extensive dataset repositories. In this setup, re…

cs.DB20232 cited

DIG: The Data Interface Grammar

Yiru Chen, Jeffery Tao, Eugene Wu

Building interactive data interfaces is hard because the design of an interface depends on the data processing needs for the underlying analysis task, yet we do not have a good rep…