1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
Rui Liu, Shuwei He, Yifan Hu +1
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in unde…
cs.CL2025★ 1 cited
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis
Rui Liu, Zhenqi Jia, Feilong Bao +1
Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlik…
cs.MM2025
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
Rui Liu, Hongyu Yuan, Haizhou Li
Unlike traditional Automatic Speech Recognition (ASR), Audio-Visual Speech Recognition (AVSR) takes audio and visual signals simultaneously to infer the transcription. Recent studi…