5 papers
SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
Inwon Kang, Kavitha Srinivas, Nandana Mihindukulasooriya +4
Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task by capturing linguistic sema…
Measuring Privacy Risks and Tradeoffs in Financial Synthetic Data Generation
Michael Zuo, Inwon Kang, Stacy Patterson +1
We explore the privacy-utility tradeoff of synthetic data generation schemes on tabular financial datasets, a domain characterized by high regulatory risk and severe class imbalanc…
Language Model Representations for Efficient Few-Shot Tabular Classification
Inwon Kang, Parikshit Ram, Yi Zhou +2
The Web is a rich source of structured data in the form of tables, from product catalogs and knowledge bases to scientific datasets. However, the heterogeneity of the structure and…
Terminators: Terms of Service Parsing and Auditing Agents
Maruf Ahmed Mridul, Inwon Kang, Oshani Seneviratne
Terms of Service (ToS) documents are often lengthy and written in complex legal language, making them difficult for users to read and understand. To address this challenge, we prop…
On Learning Representations for Tabular Data Distillation
Inwon Kang, Parikshit Ram, Yi Zhou +2
Dataset distillation generates a small set of information-rich instances from a large dataset, resulting in reduced storage requirements, privacy or copyright risks, and computatio…