108 citations · 427 across the 45 of their papers we have counts for
10 papers · 1 filter
LongFNT: Long-form Speech Recognition with Factorized Neural Transducer
Xun Gong, Yu Wu, Jinyu Li +4
Traditional automatic speech recognition~(ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is mor…
Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
Kun Wei, Long Zhou, Ziqiang Zhang +5
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity probl…
Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
Heming Wang, Yao Qian, Xiaofei Wang +6
Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a…
Streaming Multi-talker Speech Recognition with Joint Speaker Identification
Liang Lu, Naoyuki Kanda, Jinyu Li +1
In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to transcribe the audio as well as identify the speakers for downstream…
Streaming end-to-end multi-talker speech recognition
Liang Lu, Naoyuki Kanda, Jinyu Li +1
End-to-end multi-talker speech recognition is an emerging research trend in the speech community due to its vast potential in applications such as conversation and meeting transcri…
Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
Sanyuan Chen, Yu Wu, Zhuo Chen +3
With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been…