Overlearning Reveals Sensitive Attributes
arXiv:1905.11742
Abstract
"Overlearning" means that a model trained for a seemingly simple objective implicitly learns to recognize attributes and concepts that are (1) not part of the learning objective, and (2) sensitive from a privacy or bias perspective. For example, a binary gender classifier of facial images also learns to recognize races\textemdash even races that are not represented in the training data\textemdash and identities. We demonstrate overlearning in several vision and NLP models and analyze its harmful consequences. First, inference-time representations of an overlearned model reveal sensitive attributes of the input, breaking privacy protections such as model partitioning. Second, an overlearned model can be "re-purposed" for a different, privacy-violating task even in the absence of the original training data. We show that overlearning is intrinsic for some tasks and cannot be prevented by censoring unwanted attributes. Finally, we investigate where, when, and why overlearning happens during model training.
References in corpus (8)
- How transferable are features in deep neural networks?
- Similarity of Neural Network Representations Revisited
- What makes ImageNet good for transfer learning?
- Age Progression/Regression by Conditional Adversarial Autoencoder
- Learning Controllable Fair Representations
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Not Just Privacy: Improving Performance of Private Deep Learning in Mobile Cloud
- Privacy Partitioning: Protecting User Data During the Deep Learning Inference Phase
Cited by in corpus (16)
- On the Opportunities and Risks of Foundation Models
- A Survey of Privacy Attacks in Machine Learning
- Node-Level Membership Inference Attacks Against Graph Neural Networks
- ML-Doctor: Holistic Risk Assessment of Inference Attacks Against Machine Learning Models
- Honest-but-Curious Nets: Sensitive Attributes of Private Inputs Can Be Secretly Coded into the Classifiers' Outputs
- Leakage of Dataset Properties in Multi-Party Machine Learning
- Dataset Inference: Ownership Resolution in Machine Learning
- FaceLeaks: Inference Attacks against Transfer Learning Models via Black-box Queries
- Quantifying and Mitigating Privacy Risks of Contrastive Learning
- Unsupervised Information Obfuscation for Split Inference of Neural Networks
- The Connection between Out-of-Distribution Generalization and Privacy of ML Models
- Obfuscation of Images via Differential Privacy: From Facial Images to General Images
- Fair Normalizing Flows
- A Tandem Framework Balancing Privacy and Security for Voice User Interfaces
- 10 Security and Privacy Problems in Large Foundation Models
- Confidential Machine Learning on Untrusted Platforms: A Survey