115 citations · 342 across the 42 of their papers we have counts for
6 papers · 1 filter
Cross Attention Based Style Distribution for Controllable Person Image Synthesis
Xinyue Zhou, Mingyu Yin, Xinyuan Chen +3
Controllable person image synthesis task enables a wide range of applications through explicit control over body pose and appearance. In this paper, we propose a cross attention ba…
MAR: Masked Autoencoders for Efficient Action Recognition
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +5
Standard approaches for video recognition usually operate on the full input videos, which is inefficient due to the widely present spatio-temporal redundancy in videos. Recent prog…
Context-aware Proposal Network for Temporal Action Detection
Xiang Wang, Huaxin Zhang, Shiwei Zhang +3
This technical report presents our first place winning solution for temporal action detection task in CVPR-2022 AcitivityNet Challenge. The task aims to localize temporal boundarie…
Hybrid Relation Guided Set Matching for Few-shot Action Recognition
Xiang Wang, Shiwei Zhang, Zhiwu Qing +5
Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal ali…
Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +6
Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos,…
Style Transformer for Image Inversion and Editing
Xueqi Hu, Qiusheng Huang, Zhengyi Shi +4
Existing GAN inversion methods fail to provide latent codes for reliable reconstruction and flexible editing simultaneously. This paper presents a transformer-based image inversion…