Data Governance in the Age of Large-Scale Data-Driven Language Technology
arXiv:2206.03216 · doi:10.1145/3531146.3534637
Abstract
The recent emergence and adoption of Machine Learning technology, and specifically of Large Language Models, has drawn attention to the need for systematic and transparent management of language data. This work proposes an approach to global language data governance that attempts to organize data management amongst stakeholders, values, and rights. Our proposal is informed by prior work on distributed governance that accounts for human values and grounded by an international research collaboration that brings together researchers and practitioners from 60 countries. The framework we present is a multi-party international governance structure focused on language data, and incorporating technical and organizational tools needed to support its work.
32 pages: Full paper and Appendices; Association for Computing Machinery, New York, NY, USA, 2206-2222
References in corpus (12)
- Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence
- Training Compute-Optimal Large Language Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Strengthening legal protection against discrimination by algorithms and artificial intelligence
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development
- An Army of Me: Sockpuppets in Online Discussion Communities
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- Towards Standardization of Data Licenses: The Montreal Data License
- Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards
- Affirmative Algorithms: The Legal Grounds for Fairness as Awareness
- No Intruder, no Validity: Evaluation Criteria for Privacy-Preserving Text Anonymization
Cited by in corpus (7)
- Auditing large language models: a three-layered approach
- Harms from Increasingly Agentic Algorithmic Systems
- Machine Culture
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- Complex QA and language models hybrid architectures, Survey
- Evaluating the Social Impact of Generative AI Systems in Systems and Society