3 papers
cs.SD2025
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
Haorui He, Zengqiang Shang, Chaoren Wang +11
Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in…
cs.SD2025
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
Xuyuan Li, Zengqiang Shang, Hua Hua +4
Recently, neural ordinary differential equations (ODE) models trained with flow matching have achieved impressive performance on the zero-shot voice clone task. Nevertheless, postu…
cs.SD2025
Controlling your Attributes in Voice
Xuyuan Li, Zengqiang Shang. Li Wang, Pengyuan Zhang
Attribute control in generative tasks aims to modify personal attributes, such as age and gender while preserving the identity information in the source sample. Although significan…