activity
20222026
most citedNTU-NPU System for Voice Privacy 2024 Challenge

7 citations · 23 across the 42 of their papers we have counts for

collaborators
Showing cs.SDShow all

14 papers · 1 filter

cs.SD2025

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding

Dake Guo, Jixun Yao, Linhan Ma +2

Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interacti…

cs.SD2025

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Yu Pan, Yanni Hu, Yuguang Yang +5

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a no…

cs.SD2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Qing Wang, Jixun Yao, Zhaokai Sun +3

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we ai…

cs.SD2024

CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion

Yuke Li, Xinfa Zhu, Hanzhao Li +6

Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice convers…

cs.SD2024

Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching

Yu Pan, Yuguang Yang, Jixun Yao +2

Despite recent advances in zero-shot voice conversion (VC), achieving speaker similarity and naturalness comparable to ground-truth recordings remains a significant challenge. In t…

cs.SD2024

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings

Kangxiang Xia, Dake Guo, Jixun Yao +9

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneou…