Label-Only Membership Inference Attacks
arXiv:2007.14321
Abstract
Membership inference attacks are one of the simplest forms of privacy leakage for machine learning models: given a data point and model, determine whether the point was used to train the model. Existing membership inference attacks exploit models' abnormal confidence when queried on their training data. These attacks do not apply if the adversary only gets access to models' predicted labels, without a confidence measure. In this paper, we introduce label-only membership inference attacks. Instead of relying on confidence scores, our attacks evaluate the robustness of a model's predicted labels under perturbations to obtain a fine-grained membership signal. These perturbations include common data augmentations or adversarial examples. We empirically show that our label-only membership inference attacks perform on par with prior attacks that required access to model confidences. We further demonstrate that label-only attacks break multiple defenses against membership inference attacks that (implicitly or explicitly) rely on a phenomenon we call confidence masking. These defenses modify a model's confidence scores in order to thwart attacks, but leave the model's predicted labels unchanged. Our label-only attacks demonstrate that confidence-masking is not a viable defense strategy against membership inference. Finally, we investigate worst-case label-only attacks, that infer membership for a small number of outlier data points. We show that label-only attacks also match confidence-based attacks in this setting. We find that training models with differential privacy and (strong) L2 regularization are the only known defense strategies that successfully prevents all attacks. This remains true even when the differential privacy budget is too high to offer meaningful provable guarantees.
16 pages, 11 figures, 2 tables Revision 2: 19 pages, 12 figures, 3 tables. Improved text and additional experiments. Final ICML paper
References in corpus (12)
- Explaining and Harnessing Adversarial Examples
- Improved Regularization of Convolutional Neural Networks with Cutout
- The Effectiveness of Data Augmentation in Image Classification using Deep Learning
- Understanding deep learning requires rethinking generalization
- A Boundary Tilting Persepective on the Phenomenon of Adversarial Examples
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference
- Data augmentation using synthetic data for time series classification with deep residual networks
- HopSkipJumpAttack: A Query-Efficient Decision-Based Attack
- Towards Measuring Membership Privacy
- A New Defense Against Adversarial Images: Turning a Weakness into a Strength
- Defending Model Inversion and Membership Inference Attacks via Prediction Purification
Cited by in corpus (16)
- Node-Level Membership Inference Attacks Against Graph Neural Networks
- Debiasing Learning for Membership Inference Attacks Against Recommender Systems
- Dataset Inference: Ownership Resolution in Machine Learning
- Training Data Leakage Analysis in Language Models
- Knowledge Cross-Distillation for Membership Privacy
- Membership Leakage in Label-Only Exposures
- Membership Inference on Word Embedding and Beyond
- SoK: Dataset Copyright Auditing in Machine Learning Systems
- Membership Inference Attacks Against Recommender Systems
- MACE: A Flexible Framework for Membership Privacy Estimation in Generative Models
- On the Privacy Risks of Cell-Based NAS Architectures
- Quantifying and Mitigating Privacy Risks of Contrastive Learning
- LazyDP: Co-Designing Algorithm-Software for Scalable Training of Differentially Private Recommendation Models
- Correlation inference attacks against machine learning models
- TableGAN-MCA: Evaluating Membership Collisions of GAN-Synthesized Tabular Data Releasing
- On the Security Risks of AutoML