64 citations · 161 across the 7 of their papers we have counts for
7 papers
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Kunchang Li, Yali Wang, Yinan He +4
Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term v…
InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges
Guo Chen, Sen Xing, Zhe Chen +18
In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, includin…
Test-Time Personalization with a Transformer for Human Pose Estimation
Yizhuo Li, Miao Hao, Zonglin Di +2
We propose to personalize a human pose estimator given a set of test images of a person without using any manual annotations. While there is a significant advancement in human pose…
PGT: A Progressive Method for Training Models on Long Videos
Bo Pang, Gao Peng, Yizhuo Li +1
Convolutional video models have an order of magnitude larger computational complexity than their counterpart image-level models. Constrained by computational resources, there is no…
TDAF: Top-Down Attention Framework for Vision Tasks
Bo Pang, Yizhuo Li, Jiefeng Li +3
Human attention mechanisms often work in a top-down manner, yet it is not well explored in vision research. Here, we propose the Top-Down Attention Framework (TDAF) to capture top-…
HOI Analysis: Integrating and Decomposing Human-Object Interaction
Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu +2
Human-Object Interaction (HOI) consists of human, object and implicit interaction/verb. Different from previous methods that directly map pixels to HOI semantics, we propose a nove…