Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…
cs.SD2024
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
Xinyuan Qian, Jiaran Gao, Yaodan Zhang +4
Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the ma…