14 citations · 20 across the 4 of their papers we have counts for
4 papers
Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers
Yasheng Sun, Hang Zhou, Kaisiyuan Wang +7
Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facia…
Audio-Driven Co-Speech Gesture Video Generation
Xian Liu, Qianyi Wu, Hang Zhou +4
Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly…
StyleSwap: Style-Based Generator Empowers Robust Face Swapping
Zhiliang Xu, Hang Zhou, Zhibin Hong +7
Numerous attempts have been made to the task of person-agnostic face swapping given its wide applications. While existing methods mostly rely on tedious network and loss designs, t…
Incorporating Convolution Designs into Visual Transformers
Kun Yuan, Shaopeng Guo, Ziwei Liu +3
Motivated by the success of Transformers in natural language processing (NLP) tasks, there emerge some attempts (e.g., ViT and DeiT) to apply Transformers to the vision domain. How…