activity
20182025
most citedUniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

108 citations · 361 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV202293 cited

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Yi Wang, Kunchang Li, Yizhuo Li +14

The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on…

cs.CV202259 cited

UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Kunchang Li, Yali Wang, Yinan He +4

Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term v…

cs.CV202214 cited

InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Guo Chen, Sen Xing, Zhe Chen +18

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, includin…

cs.CV20221 cited

VideoPipe 2022 Challenge: Real-World Video Understanding for Urban Pipe Inspection

Yi Liu, Xuan Zhang, Ying Li +9

Video understanding is an important problem in computer vision. Currently, the well-studied task in this research is human action recognition, where the clips are manually trimmed…

cs.CV20223 cited

Cross Domain Object Detection by Target-Perceived Dual Branch Distillation

Mengzhe He, Yali Wang, Jiaxi Wu +6

Cross domain object detection is a realistic and challenging task in the wild. It suffers from performance degradation due to large shift of data distributions and lack of instance…

cs.CV20222 cited

Dual-AI: Dual-path Actor Interaction Learning for Group Activity Recognition

Mingfei Han, David Junhao Zhang, Yali Wang +4

Learning spatial-temporal relation among multiple actors is crucial for group activity recognition. Different group activities often show the diversified interactions between actor…