4 papers
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Yifan Hu, Shuwei He, Rui Liu +1
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, f…
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
Zihan Lin, Songhe Deng, Shuwei He +6
Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions im…
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
Rui Liu, Shuwei He, Yifan Hu +1
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in unde…
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
Shuwei He, Rui Liu
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality fo…