Imperfect ImaGANation: Implications of GANs Exacerbating Biases on Facial Data Augmentation and Snapchat Selfie Lenses
arXiv:2001.09528
Abstract
In this paper, we show that popular Generative Adversarial Networks (GANs) exacerbate biases along the axes of gender and skin tone when given a skewed distribution of face-shots. While practitioners celebrate synthetic data generation using GANs as an economical way to augment data for training data-hungry machine learning models, it is unclear whether they recognize the perils of such techniques when applied to real world datasets biased along latent dimensions. Specifically, we show that (1) traditional GANs further skew the distribution of a dataset consisting of engineering faculty headshots, generating minority modes less often and of worse quality and (2) image-to-image translation (conditional) GANs also exacerbate biases by lightening skin color of non-white faces and transforming female facial features to be masculine when generating faces of engineering professors. Thus, our study is meant to serve as a cautionary tale.
References in corpus (12)
- Spectral Normalization for Generative Adversarial Networks
- MesoNet: a Compact Facial Video Forgery Detection Network
- Data Augmentation Generative Adversarial Networks
- DeepFakes: a New Threat to Face Recognition? Assessment and Detection
- FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces
- The Deepfake Detection Challenge (DFDC) Preview Dataset
- Do GANs actually learn the distribution? An empirical study
- Synthetic Medical Images from Dual Generative Adversarial Networks
- Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics
- Bias and Generalization in Deep Generative Models: An Empirical Study
- Cross-modality image synthesis from unpaired data using CycleGAN: Effects of gradient consistency loss and training data size
- Mode matching in GANs through latent space learning and inversion