most citedDiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data

22 citations · 32 across the 5 of their papers we have counts for

collaborators

5 papers

cs.DB20241 cited

SketchQL Demonstration: Zero-shot Video Moment Querying with Sketches

Renzhi Wu, Pramod Chunduri, Dristi J Shah +5

In this paper, we will present SketchQL, a video database management system (VDBMS) for retrieving video moments with a sketch-based query interface. This novel interface allows us…

cs.LG20241 cited

Falcon: Fair Active Learning using Multi-armed Bandits

Ki Hyun Tae, Hantian Zhang, Jaeyoung Park +2

Biased data can lead to unfair machine learning models, highlighting the importance of embedding fairness at the beginning of data analysis, particularly during dataset curation an…

cs.DC20246 cited

Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native

Yao Lu, Song Bian, Lequn Chen +19

In this paper, we investigate the intersection of large generative AI models and cloud-native computing architectures. Recent large models such as ChatGPT, while revolutionary in t…

cs.DB202322 cited

DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data

Peng Li, Zhiyi Chen, Xu Chu +1

Data preprocessing is a crucial step in the machine learning process that transforms raw data into a more usable format for downstream ML models. However, it can be costly and time…

cs.DB20232 cited

Rethinking Similarity Search: Embracing Smarter Mechanisms over Smarter Data

Renzhi Wu, Jingfan Meng, Jie Jeff Xu +2

In this vision paper, we propose a shift in perspective for improving the effectiveness of similarity search. Rather than focusing solely on enhancing the data quality, particularl…