14 citations · 27 across the 4 of their papers we have counts for
4 papers
Audio-Driven Co-Speech Gesture Video Generation
Xian Liu, Qianyi Wu, Hang Zhou +4
Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly…
Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation
Xian Liu, Qianyi Wu, Hang Zhou +7
Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner,…
Visual Sound Localization in the Wild by Cross-Modal Interference Erasing
Xian Liu, Rui Qian, Hang Zhou +5
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios ar…
Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation
Xian Liu, Yinghao Xu, Qianyi Wu +3
Animating high-fidelity video portrait with speech audio is crucial for virtual reality and digital entertainment. While most previous studies rely on accurate explicit structural…