Lessons from Archives: Strategies for Collecting Sociocultural Data in Machine Learning
arXiv:1912.10389 · doi:10.1145/3351095.3372829
Abstract
A growing body of work shows that many problems in fairness, accountability, transparency, and ethics in machine learning systems are rooted in decisions surrounding the data collection and annotation process. In spite of its fundamental nature however, data collection remains an overlooked part of the machine learning (ML) pipeline. In this paper, we argue that a new specialization should be formed within ML that is focused on methodologies for data collection and annotation: efforts that require institutional frameworks and procedures. Specifically for sociocultural data, parallels can be drawn from archives and libraries. Archives are the longest standing communal effort to gather human information and archive scholars have already developed the language and procedures to address and discuss many challenges pertaining to data collection such as consent, power, inclusivity, transparency, and ethics & privacy. We discuss these five key approaches in document collection practices in archives that can inform data collection in sociocultural ML. By showing data collection practices from another field, we encourage ML research to be more cognizant and systematic in data collection and draw from interdisciplinary expertise.
To be published in Conference on Fairness, Accountability, and Transparency FAT* '20, January 27-30, 2020, Barcelona, Spain. ACM, New York, NY, USA, 11 pages
References in corpus (1)
Cited by in corpus (28)
- Power to the People? Opportunities and Challenges for Participatory AI
- Auditing large language models: a three-layered approach
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
- Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation
- Data Feminism for AI
- The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations
- Can Workers Meaningfully Consent to Workplace Wellbeing Technologies?
- Robots Enact Malignant Stereotypes
- The worst of both worlds: A comparative analysis of errors in learning from data in psychology and machine learning
- Computer Vision and Conflicting Values: Describing People with Automated Alt Text
- The craft and coordination of data curation: complicating "workflow" views of data science
- Social Inclusion in Curated Contexts: Insights from Museum Practices
- Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
- Skin Deep: Investigating Subjectivity in Skin Tone Annotations for Computer Vision Benchmark Datasets
- Contributing to Accessibility Datasets: Reflections on Sharing Study Data by Blind People
- Epistemic Power in AI Ethics Labor: Legitimizing Located Complaints
- Creative Writers' Attitudes on Writing as Training Data for Large Language Models
- Machine Learning Data Practices through a Data Curation Lens: An Evaluation Framework
- Sharing Practices for Datasets Related to Accessibility and Aging
- Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
- The Power of Absence: Thinking with Archival Theory in Algorithmic Design
- AccessShare: Co-designing Data Access and Sharing with Blind People
- We Haven't Gone Paperless Yet: Why the Printing Press Can Help Us Understand Data and AI
- Algorithms in the Stacks: Investigating automated, for-profit diversity audits in public libraries
- Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals
- Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
- Algorithmic Fairness Datasets: the Story so Far
- A Labeling Task Design for Supporting Algorithmic Needs: Facilitating Worker Diversity and Reducing AI Bias