34 citations · 41 across the 5 of their papers we have counts for
5 papers
AiluRus: A Scalable ViT Framework for Dense Prediction
Jin Li, Yaoming Wang, Xiaopeng Zhang +6
Vision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, when it comes to handling long token sequences,…
Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation
Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang +3
Transformers have become the primary backbone of the computer vision community due to their impressive performance. However, the unfriendly computation cost impedes their potential…
ControlVideo: Training-free Controllable Text-to-Video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temp…
Multi-modal Prompting for Low-Shot Temporal Action Localization
Chen Ju, Zeqian Li, Peisen Zhao +5
In this paper, we consider the problem of temporal action localization under low-shot (zero-shot & few-shot) scenario, with the goal of detecting and classifying the action instanc…
Active Pointly-Supervised Instance Segmentation
Chufeng Tang, Lingxi Xie, Gang Zhang +3
The requirement of expensive annotations is a major burden for training a well-performed instance segmentation model. In this paper, we present an economic active learning setting,…