activity
20182022
most citedApplying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

58 citations · 129 across the 9 of their papers we have counts for

collaborators

14 papers

cs.SD2022

Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire

Zhiyun Fan, Zhenlin Liang, Linhao Dong +6

In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to segment the audio and then transcribe each segmentation. These two s…

cs.CL2022

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

Minglun Han, Linhao Dong, Zhenlin Liang +4

Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on ph…

cs.CV202121 cited

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Jing Liu, Xinxin Zhu, Fei Liu +8

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructe…

cs.CL2021

Efficiently Fusing Pretrained Acoustic and Linguistic Encoders for Low-resource Speech Recognition

Cheng Yi, Shiyu Zhou, Bo Xu

End-to-end models have achieved impressive results on the task of automatic speech recognition (ASR). For low-resource ASR tasks, however, labeled data can hardly satisfy the deman…

cs.CL202158 cited

Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

Cheng Yi, Jianzhong Wang, Ning Cheng +2

There are several domains that own corresponding widely used feature extractors, such as ResNet, BERT, and GPT-x. These models are usually pre-trained on large amounts of unlabeled…

cs.SD202114 cited

Exploring wav2vec 2.0 on speaker verification and language identification

Zhiyun Fan, Meng Li, Shiyu Zhou +1

Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and perfor…