activity
20202022
most citedWNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition

9 citations · 18 across the 7 of their papers we have counts for

collaborators

7 papers

eess.AS20221 cited

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

Ziqian Ning, Qicong Xie, Pengcheng Zhu +5

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balancing between speaker similarity, intelligibility and expressiveness. To addr…

cs.SD20221 cited

AccentSpeech: Learning Accent from Crowd-sourced Data for Target Speaker TTS with Accents

Yongmao Zhang, Zhichao Wang, Peiji Yang +3

Learning accent from crowd-sourced data is a feasible way to achieve a target speaker TTS system that can synthesize accent speech. To this end, there are two challenging problems…

eess.AS2022

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

Yuanzhe Chen, Ming Tu, Tang Li +7

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracte…

eess.AS20227 cited

IQDUBBING: Prosody modeling based on discrete self-supervised speech representation for expressive voice conversion

Wendong Gan, Bolong Wen, Ying Yan +6

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, wh…

eess.AS2021

Enriching Source Style Transfer in Recognition-Synthesis based Non-Parallel Voice Conversion

Zhichao Wang, Xinyong Zhou, Fengyu Yang +6

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitation…

cs.SD20219 cited

WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition

Zhichao Wang, Wenwen Yang, Pan Zhou +1

Recently, attention-based encoder-decoder (AED) end-to-end (E2E) models have drawn more and more attention in the field of automatic speech recognition (ASR). AED models, however,…