activity
20172024
most citedAIM: Adapting Image Models for Efficient Video Action Recognition

62 citations · 236 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV20231 cited

DiffHand: End-to-End Hand Mesh Reconstruction via Diffusion Models

Lijun Li, Li'an Zhuo, Bang Zhang +2

Hand mesh reconstruction from the monocular image is a challenging task due to its depth ambiguity and severe occlusion, there remains a non-unique mapping between the monocular im…

cs.CV20232 cited

Part Aware Contrastive Learning for Self-Supervised Action Recognition

Yilei Hua, Wenhan Wu, Ce Zheng +4

In recent years, remarkable results have been achieved in self-supervised action recognition using skeleton sequences with contrastive learning. It has been observed that the seman…

cs.CV20236 cited

PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

Qitao Zhao, Ce Zheng, Mengyuan Liu +2

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relation…

cs.CV20232 cited

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen +1

Semi-Supervised Learning can be more beneficial for the video domain compared to images because of its higher annotation cost and dimensionality. Besides, any video understanding t…

cs.CV2023

POTTER: Pooling Attention Transformer for Efficient Human Mesh Recovery

Ce Zheng, Xianpeng Liu, Guo-Jun Qi +1

Transformer architectures have achieved SOTA performance on the human mesh recovery (HMR) from monocular images. However, the performance gain has come at the cost of substantial m…

cs.CV202362 cited

AIM: Adapting Image Models for Efficient Video Action Recognition

Taojiannan Yang, Yi Zhu, Yusheng Xie +3

Recent vision transformer based video models mostly follow the ``image pre-training then finetuning" paradigm and have achieved great success on multiple video benchmarks. However,…