activity
20212026
most citedOn Shapley Value in Data Assemblage Under Independent Utility

15 citations · 32 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CR2026

RaMark: Radioactive Watermarking for Generated Tabular Data

Xin Che, Lingyang Chu, Qiqi Zhang +3

Recent advances in generative modeling have made generated tabular data a practical solution for privacy-sensitive data sharing, where watermarking enables ownership verification.…

cs.LG2023

Serverless Federated AUPRC Optimization for Multi-Party Collaborative Imbalanced Data Mining

Xidong Wu, Zhengmian Hu, Jian Pei +1

Multi-party collaborative training, such as distributed learning and federated learning, is used to address the big data challenges. However, traditional multi-party collaborative…

cs.CR2022

DP2-Pub: Differentially Private High-Dimensional Data Publication with Invariant Post Randomization

Honglu Jiang, Haotian Yu, Xiuzhen Cheng +3

A large amount of high-dimensional and heterogeneous data appear in practical applications, which are often published to third parties for data analysis, recommendations, targeted…

cs.LG2022★ 1 cited

Knowledge-Injected Federated Learning

Zhenan Fan, Zirui Zhou, Jian Pei +4

Federated learning is an emerging technique for training models from decentralized data sets. In many applications, data owners participating in the federated learning system hold…

cs.DB2022★ 15 cited

On Shapley Value in Data Assemblage Under Independent Utility

Xuan Luo, Jian Pei, Zicun Cong +1

In many applications, an organization may want to acquire data from many data owners. Data marketplaces allow data owners to produce data assemblage needed by data buyers through c…

cs.LG2022★ 1 cited

Revealing Unfair Models by Mining Interpretable Evidence

Mohit Bajaj, Lingyang Chu, Vittorio Romaniello +5

The popularity of machine learning has increased the risk of unfair models getting deployed in high-stake applications, such as justice system, drug/vaccination design, and medical…