3 papers
cs.CV2024
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
Hannan Lu, Xiaohe Wu, Shudong Wang +5
Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Exist…
cs.CV2024
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
Wenjun Li, Shudong Wang, Dong Zhao +3
The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image fr…
cs.CV2024
Controllable Talking Face Generation by Implicit Facial Keypoints Editing
Dong Zhao, Jiaying Shi, Wenjun Li +3
Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures…