2 citations · 2 across the 2 of their papers we have counts for
4 papers
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
Zhenwei Dai, Chuan Lei, Asterios Katsifodimos +3
How to generate a large, realistic set of tables along with joinability relationships, to stress-test dataset discovery methods? Dataset discovery methods aim to automatically iden…
FeatNavigator: Automatic Feature Augmentation on Tabular Data
Jiaming Liang, Chuan Lei, Xiao Qin +4
Data-centric AI focuses on understanding and utilizing high-quality, relevant data in training machine learning (ML) models, thereby increasing the likelihood of producing accurate…
OmniMatch: Effective Self-Supervised Any-Join Discovery in Tabular Data Repositories
Christos Koutras, Jiani Zhang, Xiao Qin +5
How can we discover join relationships among columns of tabular data in a data repository? Can this be done effectively when metadata is missing? Traditional column matching works…
Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent Space
Hengrui Zhang, Jiani Zhang, Balasubramaniam Srinivasan +5
Recent advances in tabular data generation have greatly enhanced synthetic data quality. However, extending diffusion models to tabular data is challenging due to the intricately v…