13 papers
SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
Inwon Kang, Kavitha Srinivas, Nandana Mihindukulasooriya +4
Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task by capturing linguistic sema…
Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks
Liane Vogel, Kavitha Srinivas, Niharika D'Souza +3
Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic sea…
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
Liangliang Zhang, Nandana Mihindukulasooriya, Niharika S. D'Souza +4
Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data a…
Agentic Control Center for Data Product Optimization
Priyadarshini Tamilselvan, Gregory Bramble, Sola Shirai +3
Data products enable end users to gain greater insights about their data by providing supporting assets, such as example question-SQL pairs which can be answered using the data or…
Automatic Prompt Engineering with No Task Cues and No Tuning
Faisal Chowdhury, Nandana Mihindukulasooriya, Niharika S D'Souza +4
This paper presents a system for automatic prompt engineering that is much simpler in both design and application and yet as effective as the existing approaches. It requires no tu…
DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
Faisal Chowdhury, Sola Shirai, Sarthak Dash +2
A data product is created with the intention of solving a specific problem, addressing a specific business usecase or meeting a particular need, going beyond just serving data as a…