34 citations · 36 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 2 cited
AiluRus: A Scalable ViT Framework for Dense Prediction
Jin Li, Yaoming Wang, Xiaopeng Zhang +6
Vision transformers (ViTs) have emerged as a prevalent architecture for vision tasks owing to their impressive performance. However, when it comes to handling long token sequences,…
cs.CV2023★ 34 cited
ControlVideo: Training-free Controllable Text-to-Video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang +3
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temp…
cs.CV2022
SdAE: Self-distillated Masked Autoencoder
Yabo Chen, Yuchen Liu, Dongsheng Jiang +4
With the development of generative-based self-supervised learning (SSL) approaches like BeiT and MAE, how to learn good representations by masking random patches of the input image…