collaborators

6 papers

cs.LG2026

SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching

Inwon Kang, Kavitha Srinivas, Nandana Mihindukulasooriya +4

Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task by capturing linguistic sema…

cs.LG2026

Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

Liane Vogel, Kavitha Srinivas, Niharika D'Souza +3

Tabular foundation models aim to learn universal representations of tabular data that transfer across tasks and domains, enabling applications such as table retrieval, semantic sea…

cs.IR2026

DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text

Liangliang Zhang, Nandana Mihindukulasooriya, Niharika S. D'Souza +4

Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data a…

cs.AI2026

Agentic Control Center for Data Product Optimization

Priyadarshini Tamilselvan, Gregory Bramble, Sola Shirai +3

Data products enable end users to gain greater insights about their data by providing supporting assets, such as example question-SQL pairs which can be answered using the data or…

cs.DB2025

Bridging Business Intent and Data: A Benchmark for Automatic Relational Data Product Generation

Faisal Chowdhury, Sola Shirai, Sarthak Dash +2

A data product is designed to address a specific business need by transforming raw data into a curated, usable asset that delivers actionable insights. Despite practical advances i…

cs.CL2025

StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation

Satyananda Kashyap, Sola Shirai, Nandana Mihindukulasooriya +1

Extracting structured information from text, such as key-value pairs that could augment tabular data, is quite useful in many enterprise use cases. Although large language models (…