10 papers
Privacy Auditing Synthetic Data Release through Local Likelihood Attacks
Joshua Ward, Chi-Hua Wang, Guang Cheng
Auditing the privacy leakage of synthetic data is an important but unresolved problem. Existing privacy auditing frameworks for synthetic data rely on heuristics and unrealistic as…
Ensembling Membership Inference Attacks Against Tabular Generative Models
Joshua Ward, Yuxuan Yang, Chi-Hua Wang +1
Membership Inference Attacks (MIAs) have emerged as a principled framework for auditing the privacy of synthetic data generated by tabular generative models, where many diverse met…
Towards High Supervised Learning Utility Training Data Generation: Data Pruning and Column Reordering
Tung Sum Thomas Kwok, Zeyong Zhang, Chi-Hua Wang +1
Tabular data synthesis for supervised learning ('SL') model training is gaining popularity in industries such as healthcare, finance, and retail. Despite the progress made in tabul…
GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction
Tung Sum Thomas Kwok, Chi-Hua Wang, Guang Cheng
Tabular data synthesis involves not only multi-table synthesis but also generating multi-modal data (e.g., strings and categories), which enables diverse knowledge synthesis. Howev…
DEREC-SIMPRO: unlock Language Model benefits to advance Synthesis in Data Clean Room
Tung Sum Thomas Kwok, Chi-hua Wang, Guang Cheng
Data collaboration via Data Clean Room offers value but raises privacy concerns, which can be addressed through synthetic data and multi-table synthesizers. Common multi-table synt…
Data Deletion for Linear Regression with Noisy SGD
Zhangjie Xia, Chi-Hua Wang, Guang Cheng
In the current era of big data and machine learning, it's essential to find ways to shrink the size of training dataset while preserving the training performance to improve efficie…