Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning
Man Wang, Chenyang Liu, Wenjun Li +5
Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images. Despite notable progress,…
cs.CV2024
Controllable Talking Face Generation by Implicit Facial Keypoints Editing
Dong Zhao, Jiaying Shi, Wenjun Li +3
Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures…
cs.CV2024
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
Wenjun Li, Shudong Wang, Dong Zhao +3
The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image fr…