3 citations · 7 across the 6 of their papers we have counts for
5 papers · 1 filter
Semantic Mask for Transformer based End-to-End Speech Recognition
Chengyi Wang, Yu Wu, Yujiao Du +7
Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks. This approach takes advantage of t…
Advancing Acoustic-to-Word CTC Model with Attention and Mixed-Units
Amit Das, Jinyu Li, Guoli Ye +2
The acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion is a natural end-to-end (E2E) system directly targeting word as output unit. Two issue…
Developing Far-Field Speaker System Via Teacher-Student Learning
Jinyu Li, Rui Zhao, Zhuo Chen +4
In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt…
Advancing Acoustic-to-Word CTC Model
Jinyu Li, Guoli Ye, Amit Das +2
The acoustic-to-word model based on the connectionist temporal classification (CTC) criterion was shown as a natural end-to-end (E2E) model directly targeting words as output units…
Acoustic-To-Word Model Without OOV
Jinyu Li, Guoli Ye, Rui Zhao +2
Recently, the acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion was shown as a natural end-to-end model directly targeting words as output u…