24 citations · 106 across the 16 of their papers we have counts for
18 papers
Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers
Yasheng Sun, Hang Zhou, Kaisiyuan Wang +7
Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facia…
Audio-Driven Co-Speech Gesture Video Generation
Xian Liu, Qianyi Wu, Hang Zhou +4
Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly…
Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition
Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou +6
While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their pote…
StyleSwap: Style-Based Generator Empowers Robust Face Swapping
Zhiliang Xu, Hang Zhou, Zhibin Hong +7
Numerous attempts have been made to the task of person-agnostic face swapping given its wide applications. While existing methods mostly rely on tedious network and loss designs, t…
Few-Shot Head Swapping in the Wild
Changyong Shu, Hemao Wu, Hang Zhou +7
The head swapping task aims at flawlessly placing a source head onto a target body, which is of great importance to various entertainment scenarios. While face swapping has drawn m…
SeCo: Separating Unknown Musical Visual Sounds with Consistency Guidance
Xinchi Zhou, Dongzhan Zhou, Wanli Ouyang +3
Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing dataset…