most citedFreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter

1 citations · 2 across the 8 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2024

Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Linhan Ma, Xinfa Zhu, Yuanjun Lv +5

Zero-shot voice conversion (VC) aims to transform source speech into arbitrary unseen target voice while keeping the linguistic content unchanged. Recent VC methods have made signi…

cs.SD20241 cited

FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter

Yuanjun Lv, Hai Li, Ying Yan +3

Vocoders reconstruct speech waveforms from acoustic features and play a pivotal role in modern TTS systems. Frequent-domain GAN vocoders like Vocos and APNet2 have recently seen ra…

cs.SD2024

RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention

Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan +5

In real-time speech communication systems, speech signals are often degraded by multiple distortions. Recently, a two-stage Repair-and-Denoising network (RaD-Net) was proposed with…

cs.SD2024

RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement

Mingshuai Liu, Zhuangqi Chen, Xiaopeng Yan +5

This paper introduces our repairing and denoising network (RaD-Net) for the ICASSP 2024 Speech Signal Improvement (SSI) Challenge. We extend our previous framework based on a two-s…

cs.SD20231 cited

Vec-Tok Speech: speech vectorization and tokenization for neural speech generation

Xinfa Zhu, Yuanjun Lv, Yi Lei +5

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the curre…

cs.SD2023

SALT: Distinguishable Speaker Anonymization Through Latent Space Transformation

Yuanjun Lv, Jixun Yao, Peikun Chen +3

Speaker anonymization aims to conceal a speaker's identity without degrading speech quality and intelligibility. Most speaker anonymization systems disentangle the speaker represen…