Towards Fairer Datasets: Filtering and Balancing the Distribution of the People Subtree in the ImageNet Hierarchy
arXiv:1912.07726 · doi:10.1145/3351095.3375709
Abstract
Computer vision technology is being used by many but remains representative of only a few. People have reported misbehavior of computer vision models, including offensive prediction results and lower performance for underrepresented groups. Current computer vision models are typically developed using datasets consisting of manually annotated images or videos; the data and label distributions in these datasets are critical to the models' behavior. In this paper, we examine ImageNet, a large-scale ontology of images that has spurred the development of many modern computer vision methods. We consider three key factors within the "person" subtree of ImageNet that may lead to problematic behavior in downstream computer vision technology: (1) the stagnant concept vocabulary of WordNet, (2) the attempt at exhaustive illustration of all categories with images, and (3) the inequality of representation in the images within concepts. We seek to illuminate the root causes of these concerns and take the first steps to mitigate them constructively.
Accepted to FAT* 2020
Cited by in corpus (21)
- Multi-Modal Knowledge Graph Construction and Application: A Survey
- Galaxy Zoo DECaLS: Detailed Visual Morphology Measurements from Volunteers and Deep Learning for 314,000 Galaxies
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- Speciesist bias in AI -- How AI applications perpetuate discrimination and unfair outcomes against animals
- Practical Galaxy Morphology Tools from Deep Supervised Representation Learning
- WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?
- Four Years of FAccT: A Reflexive, Mixed-Methods Analysis of Research Contributions, Shortcomings, and Future Prospects
- Real Risks of Fake Data: Synthetic Data, Diversity-Washing and Consent Circumvention
- Data Representativeness in Accessibility Datasets: A Meta-Analysis
- Computer Vision and Conflicting Values: Describing People with Automated Alt Text
- A Step Toward More Inclusive People Annotations for Fairness
- It's always personal: Using Early Exits for Efficient On-Device CNN Personalisation
- Toward Operationalizing Pipeline-aware ML Fairness: A Research Agenda for Developing Practical Guidelines and Tools
- Bias and Fairness in Computer Vision Applications of the Criminal Justice System
- Cooperative Colorization: Exploring Latent Cross-Domain Priors for NIR Image Spectrum Translation
- SudokuSens: Enhancing Deep Learning Robustness for IoT Sensing Applications using a Generative Approach
- Using Positive Matching Contrastive Loss with Facial Action Units to mitigate bias in Facial Expression Recognition
- Simplicity Bias Leads to Amplified Performance Disparities
- Algorithmic Fairness Datasets: the Story so Far
- A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety
- Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization