45 citations · 66 across the 5 of their papers we have counts for
9 papers · 1 filter
StableAnimator: High-Quality Identity-Preserving Human Image Animation
Shuyuan Tu, Zhen Xing, Xintong Han +4
Current diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffus…
HiViT: Hierarchical Vision Transformer Meets Masked Image Modeling
Xiaosong Zhang, Yunjie Tian, Wei Huang +4
Recently, masked image modeling (MIM) has offered a new methodology of self-supervised pre-training of vision transformers. A key idea of efficient implementation is to discard the…
Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
Shaobo Min, Qi Dai, Hongtao Xie +3
Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual an…
Self-Supervised Learning with Swin Transformers
Zhenda Xie, Yutong Lin, Zhuliang Yao +4
We are witnessing a modeling shift from CNN to Transformers in computer vision. In this work, we present a self-supervised learning approach called MoBY, with Vision Transformers a…
Temporal Action Detection with Multi-level Supervision
Baifeng Shi, Qi Dai, Judy Hoffman +3
Training temporal action detection in videos requires large amounts of labeled data, yet such annotation is expensive to collect. Incorporating unlabeled or weakly-labeled data to…
Weakly-Supervised Action Localization by Generative Attention Modeling
Baifeng Shi, Qi Dai, Yadong Mu +1
Weakly-supervised temporal action localization is a problem of learning an action localization model with only video-level action labeling available. The general framework largely…