activity
20222024
most citedSimVTP: Simple Video Text Pre-training with Masked Autoencoders

10 citations · 11 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing

Jiangshan Wang, Yue Ma, Jiayi Guo +3

Video editing is an emerging task, in which most current methods adopt the pre-trained text-to-image (T2I) diffusion model to edit the source video in a zero-shot manner. Despite e…

cs.CV20241 cited

Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Yue Ma, Yingqing He, Hongfa Wang +8

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and t…

cs.CV2023

MagicStick: Controllable Video Editing via Control Handle Transformations

Yue Ma, Xiaodong Cun, Sen Liang +5

Text-based video editing has recently attracted considerable interest in changing the style or replacing the objects with a similar structure. Beyond this, we demonstrate that prop…

cs.CV2023

Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos

Yue Ma, Yingqing He, Xiaodong Cun +5

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absen…

cs.CV202210 cited

SimVTP: Simple Video Text Pre-training with Masked Autoencoders

Yue Ma, Tianyu Yang, Yin Shan +1

This paper presents SimVTP: a Simple Video-Text Pretraining framework via masked autoencoders. We randomly mask out the spatial-temporal tubes of input video and the word tokens of…