Augmented Datasheets for Speech Datasets and Ethical Decision-Making
arXiv:2305.04672 · doi:10.1145/3593013.3594049
Abstract
Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in building equitable and robust SLT products, especially along dimensions of language, accent, dialect, variety, and speech impairment - and the intersectionality of speech features with socioeconomic and demographic features. Furthermore, there is often a lack of oversight on the underlying training data - commonly built on massive web-crawling and/or publicly available speech - with regard to the ethics of such data collection. To encourage standardized documentation of such speech data components, we introduce an augmented datasheet for speech datasets, which can be used in addition to "Datasheets for Datasets". We then exemplify the importance of each question in our augmented datasheet based on in-depth literature reviews of speech data used in domains such as machine learning, linguistics, and health. Finally, we encourage practitioners - ranging from dataset creators to researchers - to use our augmented datasheet to better define the scope, properties, and limits of speech datasets, while also encouraging consideration of data-subject protection and user community empowerment. Ethical dataset creation is not a one-size-fits-all process, but dataset creators can use our augmented datasheet to reflexively consider the social context of related SLT applications and data sources in order to foster more inclusive SLT products downstream.
To appear in 2023 ACM Conference on Fairness, Accountability, and Transparency (FAccT '23), June 12-15, Chicago, IL, USA
References in corpus (50)
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Robust Speech Recognition via Large-Scale Weak Supervision
- Improving fairness in machine learning systems: What do industry practitioners need?
- Prediction-Based Decisions and Fairness: A Catalogue of Choices, Assumptions, and Definitions
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation
- Measurement and Fairness
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- MusicLM: Generating Music From Text
- The State of Speech in HCI: Trends, Themes and Challenges
- A Survey on Neural Speech Synthesis
- LibriMix: An Open-Source Dataset for Generalizable Speech Separation
- Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
- Leakage and the Reproducibility Crisis in ML-based Science
- Personalizing ASR for Dysarthric and Accented Speech with Limited Data
- AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
- Evaluating Voice Conversion-based Privacy Protection against Informed Attackers
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Speech Quality and Testing Framework
- Quantifying Bias in Automatic Speech Recognition
- Training Neural Speech Recognition Systems with Synthetic Speech Augmentation
- A Review of Speech-centric Trustworthy Machine Learning: Privacy, Safety, and Fairness
- ATCSpeech: a multilingual pilot-controller speech corpus from real Air Traffic Control environment
- Causal datasheet: An approximate guide to practically assess Bayesian networks in the real world
- Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents
- MediaSpeech: Multilanguage ASR Benchmark and Dataset
- ATCO2 corpus: A Large-Scale Dataset for Research on Automatic Speech Recognition and Natural Language Understanding of Air Traffic Control Communications
- Ethical Considerations for Responsible Data Curation
- Improving fairness in speaker verification via Group-adapted Fusion Network
- On-Device Personalization of Automatic Speech Recognition Models for Disordered Speech
- JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification
- Improving Gender Translation Accuracy with Filtered Self-Training
- MT-Adapted Datasheets for Datasets: Template and Repository
- Mixtures of Deep Neural Experts for Automated Speech Scoring
- Considerations for Ethical Speech Recognition Datasets
- ASR-GLUE: A New Multi-task Benchmark for ASR-Robust Natural Language Understanding
- An Ethical Highlighter for People-Centric Dataset Creation
- Subword Dictionary Learning and Segmentation Techniques for Automatic Speech Recognition in Tamil and Kannada
- Global Performance Disparities Between English-Language Accents in Automatic Speech Recognition
- FT Speech: Danish Parliament Speech Corpus
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
- Machine Learning for Stuttering Identification: Review, Challenges and Future Directions
- Understanding the Tradeoffs in Client-side Privacy for Downstream Speech Tasks
- Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions
- Kurdish (Sorani) Speech to Text: Presenting an Experimental Dataset
- ASR in German: A Detailed Error Analysis
- Indian EmoSpeech Command Dataset: A dataset for emotion based speech recognition in the wild
- Combining Unsupervised and Text Augmented Semi-Supervised Learning for Low Resourced Autoregressive Speech Recognition
- Weaving Privacy and Power: On the Privacy Practices of Labor Organizers in the U.S. Technology Industry
- Building an ASR Error Robust Spoken Virtual Patient System in a Highly Class-Imbalanced Scenario Without Speech Data
Cited by in corpus (5)
- Careless Whisper: Speech-to-Text Hallucination Harms
- Lazy Data Practices Harm Fairness Research
- Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
- From Model Performance to Claim: How a Change of Focus in Machine Learning Replicability Can Help Bridge the Responsibility Gap
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study