2 citations · 5 across the 6 of their papers we have counts for
5 papers · 1 filter
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
Jiaming Zhang, Shengming Cao, Rui Li +8
Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Binding process of the dominant Refer…
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
Kaibin Tian, Yanhua Cheng, Yi Liu +3
In recent years, text-to-video retrieval methods based on CLIP have experienced rapid development. The primary direction of evolution is to exploit the much wider gamut of visual a…
Edit As You Wish: Video Caption Editing with Multi-grained User Control
Linli Yao, Yuanmeng Zhang, Ziheng Wang +5
Automatically narrating videos in natural language complying with user requests, i.e. Controllable Video Captioning task, can help people manage massive videos with desired intenti…
Dual-Level Decoupled Transformer for Video Captioning
Yiqi Gao, Xinglin Hou, Wei Suo +4
Video captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generat…
CapOnImage: Context-driven Dense-Captioning on Image
Yiqi Gao, Xinglin Hou, Yuanmeng Zhang +3
Existing image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation. However, texts can also be…