2 citations · 4 across the 5 of their papers we have counts for
8 papers
Controlling your Attributes in Voice
Xuyuan Li, Zengqiang Shang. Li Wang, Pengyuan Zhang
Attribute control in generative tasks aims to modify personal attributes, such as age and gender while preserving the identity information in the source sample. Although significan…
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
Haorui He, Zengqiang Shang, Chaoren Wang +11
Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in…
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
Xuyuan Li, Zengqiang Shang, Hua Hua +4
Recently, neural ordinary differential equations (ODE) models trained with flow matching have achieved impressive performance on the zero-shot voice clone task. Nevertheless, postu…
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Haorui He, Zengqiang Shang, Chaoren Wang +11
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech rem…
Enhancing Spoofing Speech Detection Using Rhythm Information
Jingze Lu, Yuxiang Zhang, Wenchao Wang +2
Current spoofing speech detection systems need more convincing evidence. In this paper, the flaws of rhythm information inherent in the TTS-generated speech are analyzed to increas…
Synthetic Speech Detection Based on Temporal Consistency and Distribution of Speaker Features
Yuxiang Zhang, Zhuo Li, Jingze Lu +2
Current synthetic speech detection (SSD) methods perform well on certain datasets but still face issues of robustness and interpretability. A possible reason is that these methods…