CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
arXiv:2004.09249
Abstract
Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and further considers the problem of distant multi-microphone conversational speech diarization and recognition in everyday home environments. Speech material is the same as the previous CHiME-5 recordings except for accurate array synchronization. The material was elicited using a dinner party scenario with efforts taken to capture data that is representative of natural conversational speech. This paper provides a baseline description of the CHiME-6 challenge for both segmented multispeaker speech recognition (Track 1) and unsegmented multispeaker speech recognition (Track 2). Of note, Track 2 is the first challenge activity in the community to tackle an unsegmented multispeaker speech recognition scenario with a complete set of reproducible open source baselines providing speech enhancement, speaker diarization, and speech recognition modules.
Cited by in corpus (28)
- Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
- VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
- Earnings-21: A Practical Benchmark for ASR in the Wild
- End-to-End Far-Field Speech Recognition with Unified Dereverberation and Beamforming
- EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments
- USTC-NELSLIP System Description for DIHARD-III Challenge
- Towards a Competitive End-to-End Speech Recognition for CHiME-6 Dinner Party Transcription
- Continuous Speech Separation with Conformer
- Lhotse: a speech data representation library for the modern deep learning ecosystem
- Libri-adhoc40: A dataset collected from synchronized ad-hoc microphone arrays
- Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations
- INTERSPEECH 2021 ConferencingSpeech Challenge: Towards Far-field Multi-Channel Speech Enhancement for Video Conferencing
- Speeding Up Permutation Invariant Training for Source Separation
- Microsoft Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2020
- Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement
- Speaker activity driven neural speech extraction
- Probing Acoustic Representations for Phonetic Properties
- Efficient Integration of Multi-channel Information for Speaker-independent Speech Separation
- Scaling sparsemax based channel selection for speech recognition with ad-hoc microphone arrays
- Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
- CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription
- Quaternion Neural Networks for Multi-channel Distant Speech Recognition
- Mixture of Speaker-type PLDAs for Children's Speech Diarization
- End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection
- WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
- Exploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcription