10 papers
PiLLar: Matching for Pivot Table Schema via LLM-guided Monte-Carlo Tree Search
Yunjun Gao, Chuangyu Ouyang, Congcong Ge +1
Pivot tables are ubiquitous in data lakes of modern data ecosystems, making accurate schema matching over pivot tables a key prerequisite for data integration. In this paper, we fo…
All-in-one Graph-based Indexing for Hybrid Search on GPUs
Zhonggen Li, Yougen Li, Yifan Zhu +3
Hybrid search has emerged as a promising paradigm that combines lexical and semantic retrieval, enhancing accuracy for applications such as recommendations, information retrieval,…
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
Congcong Ge, Yachuan Liu, Yixuan Tang +3
In commercial systems, a pervasive requirement for automatic data preparation (ADP) is to transfer relational data from disparate sources to targets with standardized schema specif…
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
Zhonggen Li, Xiangyu Ke, Yifan Zhu +3
Approximate nearest neighbor search (ANNS) in high-dimensional vector spaces has a wide range of real-world applications. Numerous methods have been proposed to handle ANNS efficie…
Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
Mengzhao Wang, Boyu Tan, Yunjun Gao +5
Hybrid search, the integration of lexical and semantic retrieval, has become a cornerstone of modern information retrieval systems, driven by demanding applications like Retrieval-…
OneDB: A Distributed Multi-Metric Data Similarity Search System
Tang Qian, Yifan Zhu, Lu Chen +5
Increasingly massive volumes of multi-modal data are being accumulated in many {real world} settings, including in health care and e-commerce. This development calls for effective…