20 citations · 20 across the 1 of their papers we have counts for
5 papers · 1 filter
Inception Transformer
Chenyang Si, Weihao Yu, Pan Zhou +3
Recent studies show that Transformer has strong capability of building long-range dependencies, yet is incompetent in capturing high frequencies that predominantly convey local inf…
Mugs: A Multi-Granular Self-Supervised Learning Framework
Pan Zhou, Yichen Zhou, Chenyang Si +3
In self-supervised learning, multi-granular features are heavily desired though rarely investigated, as different downstream tasks (e.g., general and fine-grained classification) o…
Towards Accurate Human Pose Estimation in Videos of Crowded Scenes
Li Yuan, Shuning Chang, Xuecheng Nie +5
Video-based human pose estimation in crowded scenes is a challenging problem due to occlusion, motion blur, scale variation and viewpoint change, etc. Prior approaches always fail…
A Simple Baseline for Pose Tracking in Videos of Crowded Scenes
Li Yuan, Shuning Chang, Ziyuan Huang +6
This paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events\cite{lin2020human}; specifically, here we focus on Track3: Crowd Po…
Toward Accurate Person-level Action Recognition in Videos of Crowded Scenes
Li Yuan, Yichen Zhou, Shuning Chang +6
Detecting and recognizing human action in videos with crowded scenes is a challenging problem due to the complex environment and diversity events. Prior works always fail to deal w…