22 citations · 43 across the 7 of their papers we have counts for
5 papers · 1 filter
VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
Ziyi Ni, Minglun Han, Feilong Chen +4
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these wor…
A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
Linhao Dong, Cheng Yi, Jianzong Wang +4
End-to-end models are gaining wider attention in the field of automatic speech recognition (ASR). One of their advantages is the simplicity of building that directly recognizes the…
Multilingual End-to-End Speech Recognition with A Single Transformer on Low-Resource Languages
Shiyu Zhou, Shuang Xu, Bo Xu
Sequence-to-sequence attention-based models integrate an acoustic, pronunciation and language model into a single neural network, which make them very suitable for multilingual aut…
A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese
Shiyu Zhou, Linhao Dong, Shuang Xu +1
The choice of modeling units is critical to automatic speech recognition (ASR) tasks. Conventional ASR systems typically choose context-dependent states (CD-states) or context-depe…
Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese
Shiyu Zhou, Linhao Dong, Shuang Xu +1
Sequence-to-sequence attention-based models have recently shown very promising results on automatic speech recognition (ASR) tasks, which integrate an acoustic, pronunciation and l…