collaborators

6 papers

cs.SD2025

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Haorui He, Zengqiang Shang, Chaoren Wang +11

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in…

eess.AS2025

Over-the-Air Adversarial Attack Detection: from Datasets to Defenses

Li Wang, Xiaoyan Lei, Haorui He +3

Automatic Speaker Verification (ASV) systems can be used for voice-enabled applications for identity verification. However, recent studies have exposed these systems' vulnerabiliti…

cs.SD2025

Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning

Haorui He, Yuchen Song, Yuancheng Wang +6

The effectiveness of one-shot voice conversion (VC) decreases in real-world scenarios where reference speeches, which are often sourced from the internet, contain various disturban…

cs.SD2025

SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset

Yicheng Gu, Chaoren Wang, Junan Zhang +4

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singin…

cs.SD2025

Overview of the Amphion Toolkit (v0.2)

Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…

eess.AS2024

Debatts: Zero-Shot Debating Text-to-Speech Synthesis

Yiqiao Huang, Yuancheng Wang, Jiaqi Li +4

In debating, rebuttal is one of the most critical stages, where a speaker addresses the arguments presented by the opposing side. During this process, the speaker synthesizes their…