activity
20182023
most citedX-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

22 citations · 43 across the 7 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2023

VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

Ziyi Ni, Minglun Han, Feilong Chen +4

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these wor…

eess.AS2020★ 16 cited

A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition

Linhao Dong, Cheng Yi, Jianzong Wang +4

End-to-end models are gaining wider attention in the field of automatic speech recognition (ASR). One of their advantages is the simplicity of building that directly recognizes the…

eess.AS2018

Multilingual End-to-End Speech Recognition with A Single Transformer on Low-Resource Languages

Shiyu Zhou, Shuang Xu, Bo Xu

Sequence-to-sequence attention-based models integrate an acoustic, pronunciation and language model into a single neural network, which make them very suitable for multilingual aut…

eess.AS2018

A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

Shiyu Zhou, Linhao Dong, Shuang Xu +1

The choice of modeling units is critical to automatic speech recognition (ASR) tasks. Conventional ASR systems typically choose context-dependent states (CD-states) or context-depe…

eess.AS2018

Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese

Shiyu Zhou, Linhao Dong, Shuang Xu +1

Sequence-to-sequence attention-based models have recently shown very promising results on automatic speech recognition (ASR) tasks, which integrate an acoustic, pronunciation and l…