Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning
arXiv:1702.07464
Abstract
Deep Learning has recently become hugely popular in machine learning, providing significant improvements in classification accuracy in the presence of highly-structured and large databases. Researchers have also considered privacy implications of deep learning. Models are typically trained in a centralized manner with all the data being processed by the same training algorithm. If the data is a collection of users' private data, including habits, personal pictures, geographical positions, interests, and more, the centralized server will have access to sensitive information that could potentially be mishandled. To tackle this problem, collaborative deep learning models have recently been proposed where parties locally train their deep learning structures and only share a subset of the parameters in the attempt to keep their respective training sets private. Parameters can also be obfuscated via differential privacy (DP) to make information extraction even more challenging, as proposed by Shokri and Shmatikov at CCS'15. Unfortunately, we show that any privacy-preserving collaborative deep learning is susceptible to a powerful attack that we devise in this paper. In particular, we show that a distributed, federated, or decentralized deep learning approach is fundamentally broken and does not protect the training sets of honest participants. The attack we developed exploits the real-time nature of the learning process that allows the adversary to train a Generative Adversarial Network (GAN) that generates prototypical samples of the targeted training set that was meant to be private (the samples generated by the GAN are intended to come from the same distribution as the training data). Interestingly, we show that record-level DP applied to the shared parameters of the model, as suggested in previous work, is ineffective (i.e., record-level DP is not designed to address our attack).
ACM CCS'17, 16 pages, 18 figures
References in corpus (10)
- Deep Learning in Neural Networks: An Overview
- Natural Language Processing (almost) from Scratch
- WaveNet: A Generative Model for Raw Audio
- Stealing Machine Learning Models via Prediction APIs
- Invertible Conditional GANs for image editing
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Crypto-Nets: Neural Networks over Encrypted Data
- On distinguishability criteria for estimating generative models
- Defeating Image Obfuscation with Deep Learning
- To Drop or Not to Drop: Robustness, Consistency and Differential Privacy Properties of Dropout
Cited by in corpus (57)
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated Learning
- Federated Machine Learning: Concept and Applications
- Differentially Private Generative Adversarial Network
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative Models
- DarkneTZ: Towards Model Privacy at the Edge using Trusted Execution Environments
- Pervasive AI for IoT applications: A Survey on Resource-efficient Distributed Artificial Intelligence
- Understanding Membership Inferences on Well-Generalized Learning Models
- Deep Learning in Mobile and Wireless Networking: A Survey
- Federated Continual Learning via Knowledge Fusion: A Survey
- GRNN: Generative Regression Neural Network -- A Data Leakage Attack for Federated Learning
- Privacy-Preserving Deep Learning via Weight Transmission
- Analyzing Information Leakage of Updates to Natural Language Models
- PPGAN: Privacy-preserving Generative Adversarial Network
- FedCG: Leverage Conditional GAN for Protecting Privacy and Maintaining Competitive Performance in Federated Learning
- Free-riders in Federated Learning: Attacks and Defenses
- Privacy-preserving Machine Learning through Data Obfuscation
- On Lightweight Privacy-Preserving Collaborative Learning for IoT Objects
- Machine Learning in NextG Networks via Generative Adversarial Networks
- AutoGAN-based Dimension Reduction for Privacy Preservation
- Beyond Inferring Class Representatives: User-Level Privacy Leakage From Federated Learning
- Challenges of Privacy-Preserving Machine Learning in IoT
- secureTF: A Secure TensorFlow Framework
- Over-the-Air Computing for Wireless Data Aggregation in Massive IoT
- Deep Leakage from Gradients
- Adversarial Neural Network Inversion via Auxiliary Knowledge Alignment
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- Private Model Compression via Knowledge Distillation
- Assessing differentially private deep learning with Membership Inference
- DeepObfuscation: Securing the Structure of Convolutional Neural Networks via Knowledge Distillation
- Group privacy for personalized federated learning
- LOGAN: Membership Inference Attacks Against Generative Models
- User-Level Privacy-Preserving Federated Learning: Analysis and Performance Optimization
- Privacy-Preserving Distributed Expectation Maximization for Gaussian Mixture Model using Subspace Perturbation
- Key Protected Classification for Collaborative Learning
- The More, the Better? A Study on Collaborative Machine Learning for DGA Detection
- Proof of Learning (PoLe): Empowering Machine Learning with Consensus Building on Blockchains
- Deep Learning in Information Security
- Privacy-preserving Federated Bayesian Learning of a Generative Model for Imbalanced Classification of Clinical Data
- Deep Learning Towards Mobile Applications
- A Fusion-Denoising Attack on InstaHide with Data Augmentation
- Classifying the classifier: dissecting the weight space of neural networks
- Differential Privacy for Growing Databases
- Federated learning with differential privacy and an untrusted aggregator
- Source Inference Attacks in Federated Learning
- Privacy-preserving Collaborative Learning with Automatic Transformation Search
- Learning to Succeed while Teaching to Fail: Privacy in Closed Machine Learning Systems
- Reaching Data Confidentiality and Model Accountability on the CalTrain
- Federated machine learning with Anonymous Random Hybridization (FeARH) on medical records
- Robust Membership Encoding: Inference Attacks and Copyright Protection for Deep Learning
- Towards Edge-assisted Internet of Things: From Security and Efficiency Perspectives
- Private Hierarchical Clustering and Efficient Approximation
- Differentially Private Distributed Learning for Language Modeling Tasks
- Out-of-Sample Testing for GANs
- Confidential Machine Learning on Untrusted Platforms: A Survey
- Learning to Collaborate for User-Controlled Privacy
- Distributed Layer-Partitioned Training for Privacy-Preserved Deep Learning
- Corella: A Private Multi Server Learning Approach based on Correlated Queries