3 papers
cs.AI2025
HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
Ziyu Zhang, Hanzhao Li, Jingbin Hu +2
Controllable speech synthesis refers to the precise control of speaking style by manipulating specific prosodic and paralinguistic attributes, such as gender, volume, speech rate,…
eess.AS2025
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
Xinfa Zhu, Wenjie Tian, Xinsheng Wang +4
Text-to-Audio (TTA) generation is an emerging area within AI-generated content (AIGC), where audio is created from natural language descriptions. Despite growing interest, developi…
eess.AS2025
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
Hanzhao Li, Yuke Li, Xinsheng Wang +4
Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user ne…