33 citations · 43 across the 17 of their papers we have counts for
Showing cs.SDShow all
3 papers · 1 filter
cs.SD2024
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
Sen Liu, Yiwei Guo, Xie Chen +1
While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of arti…
cs.SD2024
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
Xuenan Xu, Xiaohang Xu, Zeyu Xie +3
Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events.…
cs.SD2024
Enhancing Audio Generation Diversity with Visual Information
Zeyu Xie, Baihan Li, Xuenan Xu +2
Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited re…