Publications (8)
Privacy Auditing Synthetic Data Release through Local Likelihood Attacks
Joshua Ward, Chi-Hua Wang, Guang Cheng
Auditing the privacy leakage of synthetic data is an important but unresolved problem. Existing privacy auditing frameworks for synthetic data rely on heuristics and unrealistic as…
Risk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation
Jessup Byun, Xiaofeng Lin, Joshua Ward +1
Synthetic tabular data is essential for machine learning workflows, especially for expanding small or imbalanced datasets and enabling privacy-preserving data sharing. However, sta…
Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting
Joshua Ward, Chi-Hua Wang, Guang Cheng
Synthetic tabular data has gained attention for enabling privacy-preserving data sharing. While substantial progress has been made in single-table synthetic generation where data a…
Ensembling Membership Inference Attacks Against Tabular Generative Models
Joshua Ward, Yuxuan Yang, Chi-Hua Wang +1
Membership Inference Attacks (MIAs) have emerged as a principled framework for auditing the privacy of synthetic data generated by tabular generative models, where many diverse met…
FairRR: Pre-Processing for Group Fairness through Randomized Response
Xianli Zeng, Joshua Ward, Guang Cheng
The increasing usage of machine learning models in consequential decision-making processes has spurred research into the fairness of these systems. While significant work has been…
Synth-MIA: A Testbed for Auditing Privacy Leakage in Tabular Data Synthesis
Joshua Ward, Xiaofeng Lin, Chi-Hua Wang +1
Tabular Generative Models are often argued to preserve privacy by creating synthetic datasets that resemble training data. However, auditing their empirical privacy remains challen…
Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models
Joshua Ward, Chi-Hua Wang, Guang Cheng
The promise of tabular generative models is to produce realistic synthetic data that can be shared and safely used without dangerous leakage of information from the training set. I…
When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
Joshua Ward, Bochao Gu, Chi-Hua Wang +1
Large Language Models (LLMs) have recently demonstrated remarkable performance in generating high-quality tabular synthetic data. In practice, two primary approaches have emerged f…