3 papers
cs.SD2023
Mask-CTC-based Encoder Pre-training for Streaming End-to-End Speech Recognition
Huaibo Zhao, Yosuke Higuchi, Yusuke Kida +2
Achieving high accuracy with low latency has always been a challenge in streaming end-to-end automatic speech recognition (ASR) systems. By attending to more future contexts, a str…
cs.SD2022
Conversation-oriented ASR with multi-look-ahead CBS architecture
Huaibo Zhao, Shinya Fujie, Tetsuji Ogawa +3
During conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key…
cs.SD2021
An Investigation of Enhancing CTC Model for Triggered Attention-based Streaming ASR
Huaibo Zhao, Yosuke Higuchi, Tetsuji Ogawa +1
In the present paper, an attempt is made to combine Mask-CTC and the triggered attention mechanism to construct a streaming end-to-end automatic speech recognition (ASR) system tha…