activity
20182023
most citedPose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

24 citations · 107 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

19 papers · 1 filter

cs.CV2023

ReliTalk: Relightable Talking Portrait Generation from a Single Video

Haonan Qiu, Zhaoxi Chen, Yuming Jiang +5

Recent years have witnessed great progress in creating vivid audio-driven portraits from monocular videos. However, how to seamlessly adapt the created video avatars to other scena…

cs.CV20231 cited

StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator

Jiazhi Guan, Zhanwang Zhang, Hang Zhou +8

Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous…

cs.CV2023

Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement

Jiaxiang Tang, Hang Zhou, Xiaokang Chen +4

Neural Radiance Fields (NeRF) have constituted a remarkable breakthrough in image-based 3D reconstruction. However, their implicit volumetric representations differ significantly f…

cs.CV2022

Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers

Yasheng Sun, Hang Zhou, Kaisiyuan Wang +7

Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facia…

cs.CV202214 cited

Audio-Driven Co-Speech Gesture Video Generation

Xian Liu, Qianyi Wu, Hang Zhou +4

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly…

cs.CV202217 cited

Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou +6

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their pote…