Data and its (dis)contents: A survey of dataset development and use in machine learning research
arXiv:2012.05345 · doi:10.1016/j.patter.2021.100336
Abstract
Datasets have played a foundational role in the advancement of machine learning research. They form the basis for the models we design and deploy, as well as our primary medium for benchmarking and evaluation. Furthermore, the ways in which we collect, construct and share these datasets inform the kinds of problems the field pursues and the methods explored in algorithm development. However, recent work from a breadth of perspectives has revealed the limitations of predominant practices in dataset collection and use. In this paper, we survey the many concerns raised about the way we collect and use data in machine learning and advocate that a more cautious and thorough understanding of data is necessary to address several of the practical and ethical issues of the field.
References in corpus (13)
- Shortcut Learning in Deep Neural Networks
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Do ImageNet Classifiers Generalize to ImageNet?
- Directions in Abusive Language Training Data: Garbage In, Garbage Out
- Adversarial Filters of Dataset Biases
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's Law
- Bringing the People Back In: Contesting Benchmark Machine Learning Datasets
- Large image datasets: A pyrrhic win for computer vision?
- Towards Standardization of Data Licenses: The Montreal Data License
- Robustness to Spurious Correlations via Human Annotations
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- Utility is in the Eye of the User: A Critique of NLP Leaderboards
- Measuring Social Biases of Crowd Workers using Counterfactual Queries
Cited by in corpus (14)
- On the Opportunities and Risks of Foundation Models
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- Toward a Perspectivist Turn in Ground Truthing for Predictive Computing
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
- Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
- Societal Biases in Language Generation: Progress and Challenges
- The Limits of Global Inclusion in AI Development
- Generative Art Using Neural Visual Grammars and Dual Encoders
- Retiring Adult: New Datasets for Fair Machine Learning
- The Privatization of AI Research(-ers): Causes and Potential Consequences -- From university-industry interaction to public research brain-drain?
- Exploring Data Pipelines through the Process Lens: a Reference Model forComputer Vision
- Building Legal Datasets