activity
20182023
most citedVoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

50 citations · 209 across the 23 of their papers we have counts for

collaborators
Showing 2022Show all

7 papers · 1 filter

eess.AS2022

Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network

Dongya Jia, Qiao Tian, Kainan Peng +5

The goal of accent conversion (AC) is to convert the accent of speech into the target accent while preserving the content and speaker identity. AC enables a variety of applications…

eess.AS2022

Delivering Speaking Style in Low-resource Voice Conversion with Multi-factor Constraints

Zhichao Wang, Xinsheng Wang, Lei Xie +3

Conveying the linguistic content and maintaining the source speech's speaking style, such as intonation and emotion, is essential in voice conversion (VC). However, in a low-resour…

eess.AS2022

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

Yuanzhe Chen, Ming Tu, Tang Li +7

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracte…

cs.SD2022★ 5 cited

Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech

Zhengxi Liu, Qiao Tian, Chenxu Hu +5

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms direct…

eess.AS2022★ 50 cited

VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

Haohe Liu, Xubo Liu, Qiuqiang Kong +5

Speech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, sp…

cs.SD2022★ 1 cited

NeuFA: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention Mechanism

Jingbei Li, Yi Meng, Zhiyong Wu +4

Although deep learning and end-to-end models have been widely used and shown superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art…