123 citations · 223 across the 12 of their papers we have counts for
5 papers · 1 filter
On Speaker Attribution with SURT
Desh Raj, Matthew Wiesner, Matthew Maciejewski +3
The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in…
Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition
Dongji Gao, Hainan Xu, Desh Raj +3
Training automatic speech recognition (ASR) systems requires large amounts of well-curated paired data. However, human annotators usually perform "non-verbatim" transcription, whic…
Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
Han Zhu, Dongji Gao, Gaofeng Cheng +3
When labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, p…
Blank-regularized CTC for Frame Skipping in Neural Transducer
Yifan Yang, Xiaoyu Yang, Liyong Guo +6
Neural Transducer and connectionist temporal classification (CTC) are popular end-to-end automatic speech recognition systems. Due to their frame-synchronous design, blank symbols…
Delay-penalized CTC implemented based on Finite State Transducer
Zengwei Yao, Wei Kang, Fangjun Kuang +5
Connectionist Temporal Classification (CTC) suffers from the latency problem when applied to streaming models. We argue that in CTC lattice, the alignments that can access more fut…