10 citations · 11 across the 2 of their papers we have counts for
5 papers · 1 filter
COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
Jiangshan Wang, Yue Ma, Jiayi Guo +3
Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite e…
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
Yue Ma, Yingqing He, Hongfa Wang +8
Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and t…
MagicStick: Controllable Video Editing via Control Handle Transformations
Yue Ma, Xiaodong Cun, Sen Liang +5
Text-based video editing has recently attracted considerable interest in changing the style or replacing the objects with a similar structure. Beyond this, we demonstrate that prop…
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
Yue Ma, Yingqing He, Xiaodong Cun +5
Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absen…
SimVTP: Simple Video Text Pre-training with Masked Autoencoders
Yue Ma, Tianyu Yang, Yin Shan +1
This paper presents SimVTP: a Simple Video-Text Pretraining framework via masked autoencoders. We randomly mask out the spatial-temporal tubes of input video and the word tokens of…