3 citations · 3 across the 3 of their papers we have counts for
3 papers
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
Chak-Fai Li, William Hartmann, Matthew Snover
Transcribing the speech of multiple overlapping speakers typically requires separating the audio into multiple streams and recognizing each one independently. More recent work join…
Training Autoregressive Speech Recognition Models with Limited in-domain Supervision
Chak-Fai Li, Francis Keith, William Hartmann +1
Advances in self-supervised learning have significantly reduced the amount of transcribed audio required for training. However, the majority of work in this area is focused on read…
Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts
Chak-Fai Li, Francis Keith, William Hartmann +2
Sequence-to-sequence (seq2seq) models are competitive with hybrid models for automatic speech recognition (ASR) tasks when large amounts of training data are available. However, da…