Feature Inference Attack on Model Predictions in Vertical Federated Learning
arXiv:2010.10152 · doi:10.1109/ICDE51399.2021.00023
Abstract
Federated learning (FL) is an emerging paradigm for facilitating multiple organizations' data collaboration without revealing their private data to each other. Recently, vertical FL, where the participating organizations hold the same set of samples but with disjoint features and only one organization owns the labels, has received increased attention. This paper presents several feature inference attack methods to investigate the potential privacy leakages in the model prediction stage of vertical FL. The attack methods consider the most stringent setting that the adversary controls only the trained vertical FL model and the model predictions, relying on no background information. We first propose two specific attacks on the logistic regression (LR) and decision tree (DT) models, according to individual prediction output. We further design a general attack method based on multiple prediction outputs accumulated by the adversary to handle complex models, such as neural networks (NN) and random forest (RF) models. Experimental evaluations demonstrate the effectiveness of the proposed attacks and highlight the need for designing private mechanisms to protect the prediction outputs in vertical FL.
Accepted at the IEEE 37th International Conference on Data Engineering (ICDE 2021); 15 pages
References in corpus (2)
Cited by in corpus (8)
- Vertical Federated Learning: Concepts, Advances and Challenges
- Survey on Federated Learning Threats: concepts, taxonomy on attacks and defences, experimental study and challenges
- Advances in Robust Federated Learning: A Survey with Heterogeneity Considerations
- Residue-based Label Protection Mechanisms in Vertical Logistic Regression
- Exploring Privacy and Fairness Risks in Sharing Diffusion Models: An Adversarial Perspective
- ADI: Adversarial Dominating Inputs in Vertical Federated Learning Systems
- An Information-Theoretic Analysis of The Cost of Decentralization for Learning and Inference Under Privacy Constraints
- MedLeak: Multimodal Medical Data Leakage in Secure Federated Learning with Crafted Models