9 citations · 11 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Dynamic and Compressive Adaptation of Transformers From Images to Videos
Guozhen Zhang, Jingyu Liu, Shengming Cao +4
Recently, the remarkable success of pre-trained Vision Transformers (ViTs) from image-text matching has sparked an interest in image-to-video adaptation. However, most current appr…
cs.CV2024★ 2 cited
StableDrag: Stable Dragging for Point-based Image Editing
Yutao Cui, Xiaotong Zhao, Guozhen Zhang +3
Point-based image editing has attracted remarkable attention since the emergence of DragGAN. Recently, DragDiffusion further pushes forward the generative quality via adapting this…
cs.CV2022★ 9 cited
TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval
Yuqi Liu, Pengfei Xiong, Luhui Xu +2
Text-Video retrieval is a task of great practical value and has received increasing attention, among which learning spatial-temporal video representation is one of the research hot…