1 paper · 1 filter
Shuwei He, Rui Liu
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality fo…