14 citations · 73 across the 55 of their papers we have counts for
3 papers · 1 filter
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
Zhenwei Dai, Chuan Lei, Asterios Katsifodimos +3
How to generate a large, realistic set of tables along with joinability relationships, to stress-test dataset discovery methods? Dataset discovery methods aim to automatically iden…
CoddLLM: Empowering Large Language Models for Data Analytics
Jiani Zhang, Hengrui Zhang, Rishav Chakravarti +6
Large Language Models (LLMs) have the potential to revolutionize data analytics by simplifying tasks such as data discovery and SQL query synthesis through natural language interac…
FeatNavigator: Automatic Feature Augmentation on Tabular Data
Jiaming Liang, Chuan Lei, Xiao Qin +4
Data-centric AI focuses on understanding and utilizing high-quality, relevant data in training machine learning (ML) models, thereby increasing the likelihood of producing accurate…