4 citations · 10 across the 6 of their papers we have counts for
6 papers
Dynamic and Compressive Adaptation of Transformers From Images to Videos
Guozhen Zhang, Jingyu Liu, Shengming Cao +4
Recently, the remarkable success of pre-trained Vision Transformers (ViTs) from image-text matching has sparked an interest in image-to-video adaptation. However, most current appr…
Efficient Test-Time Prompt Tuning for Vision-Language Models
Yuhan Zhu, Guozhen Zhang, Chen Xu +4
Vision-language models have showcased impressive zero-shot classification capabilities when equipped with suitable text prompts. Previous studies have shown the effectiveness of te…
StableDrag: Stable Dragging for Point-based Image Editing
Yutao Cui, Xiaotong Zhao, Guozhen Zhang +3
Point-based image editing has attracted remarkable attention since the emergence of DragGAN. Recently, DragDiffusion further pushes forward the generative quality via adapting this…
MGMAE: Motion Guided Masking for Video Masked Autoencoding
Bingkun Huang, Zhiyu Zhao, Guozhen Zhang +2
Masked autoencoding has shown excellent performance on self-supervised video representation learning. Temporal redundancy has led to a high masking ratio and customized masking str…
DPL: Decoupled Prompt Learning for Vision-Language Models
Chen Xu, Yuhan Zhu, Guozhen Zhang +5
Prompt learning has emerged as an efficient and effective approach for transferring foundational Vision-Language Models (e.g., CLIP) to downstream tasks. However, current methods t…
Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation
Guozhen Zhang, Yuhan Zhu, Haonan Wang +3
Effectively extracting inter-frame motion and appearance information is important for video frame interpolation (VFI). Previous works either extract both types of information in a…