activity
20182022
most citedApplying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

58 citations · 80 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CL2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

Minglun Han, Linhao Dong, Zhenlin Liang +4

Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on ph…

cs.CL20216 cited

MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition

Linghui Meng, Jin Xu, Xu Tan +3

In this paper, we propose MixSpeech, a simple yet effective data augmentation method based on mixup for automatic speech recognition (ASR). MixSpeech trains an ASR model by taking…

cs.CL2021

Efficiently Fusing Pretrained Acoustic and Linguistic Encoders for Low-resource Speech Recognition

Cheng Yi, Shiyu Zhou, Bo Xu

End-to-end models have achieved impressive results on the task of automatic speech recognition (ASR). For low-resource ASR tasks, however, labeled data can hardly satisfy the deman…

cs.CL202158 cited

Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

Cheng Yi, Jianzhong Wang, Ning Cheng +2

There are several domains that own corresponding widely used feature extractors, such as ResNet, BERT, and GPT-x. These models are usually pre-trained on large amounts of unlabeled…

cs.CL2020

"Listen, Understand and Translate": Triple Supervision Decouples End-to-end Speech-to-text Translation

Qianqian Dong, Rong Ye, Mingxuan Wang +4

An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel cor…

eess.AS202016 cited

A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition

Linhao Dong, Cheng Yi, Jianzong Wang +4

End-to-end models are gaining wider attention in the field of automatic speech recognition (ASR). One of their advantages is the simplicity of building that directly recognizes the…