5 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
Zhe Niu, Brian Mak
Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show…