1 paper · 1 filter
Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4
Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the hierarchical nature of speech,…