3 papers
cs.CL2021
Spoken Term Detection Methods for Sparse Transcription in Very Low-resource Settings
Éric Le Ferrand, Steven Bird, Laurent Besacier
We investigate the efficiency of two very different spoken term detection approaches for transcription when the available data is insufficient to train a robust ASR system. This wo…
cs.CL2020
Enabling Interactive Transcription in an Indigenous Community
Éric Le Ferrand, Steven Bird, Laurent Besacier
We propose a novel transcription workflow which combines spoken term detection and human-in-the-loop, together with a pilot experiment. This work is grounded in an almost zero-reso…
cs.CL2019
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible
Marcely Zanon Boito, William N. Havard, Mahault Garnerin +2
The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to b…