activity
20202026
most citedMulti-head Monotonic Chunkwise Attention For Online Speech Recognition

13 citations · 20 across the 19 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2025

PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

Songjun Cao, Qinghua Wu, Jie Chen +2

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system com…

eess.AS20231 cited

DistillW2V2: A Small and Streaming Wav2vec 2.0 Based ASR Model

Yanzhe Fu, Yueteng Kang, Songjun Cao +1

Wav2vec 2.0 (W2V2) has shown impressive performance in automatic speech recognition (ASR). However, the large model size and the non-streaming architecture make it hard to be used…

eess.AS2022

A practical framework for multi-domain speech recognition and an instance sampling method to neural language modeling

Yike Zhang, Xiaobing Feng, Yi Liu +2

Automatic speech recognition (ASR) systems used on smart phones or vehicles are usually required to process speech queries from very different domains. In such situations, a vanill…

eess.AS20214 cited

Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-supervised Learning

Keqi Deng, Songjun Cao, Long Ma

Recently, self-supervised pre-training has gained success in automatic speech recognition (ASR). However, considering the difference between speech accents in real scenarios, how t…

eess.AS2021

Improving Streaming Transformer Based ASR Under a Framework of Self-supervised Learning

Songjun Cao, Yueteng Kang, Yanzhe Fu +4

Recently self-supervised learning has emerged as an effective approach to improve the performance of automatic speech recognition (ASR). Under such a framework, the neural network…

eess.AS2021

Improving Speech Recognition Accuracy of Local POI Using Geographical Models

Songjun Cao, Yike Zhang, Xiaobing Feng +1

Nowadays voice search for points of interest (POI) is becoming increasingly popular. However, speech recognition for local POI has remained to be a challenge due to multi-dialect a…