385 citations · 1.5k across the 33 of their papers we have counts for
42 papers · 1 filter
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Yi Wang, Kunchang Li, Yizhuo Li +14
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on…
VLG: General Video Recognition with Web Textual Knowledge
Jintao Lin, Zhaoyang Liu, Wenhai Wang +2
Video recognition in an open and dynamic world is quite challenging, as we need to handle different settings such as close-set, long-tail, few-shot and open-set. By leveraging sema…
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Kunchang Li, Yali Wang, Yinan He +4
Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term v…
AdaMixer: A Fast-Converging Query-Based Object Detector
Ziteng Gao, Limin Wang, Bing Han +1
Traditional object detectors employ the dense paradigm of scanning over locations and scales in an image. The recent query-based object detectors break this convention by decoding…
Task-specific Inconsistency Alignment for Domain Adaptive Object Detection
Liang Zhao, Limin Wang
Detectors trained with massive labeled data often exhibit dramatic performance degradation in some particular scenarios with data distribution gap. To alleviate this problem of dom…
MixFormer: End-to-End Tracking with Iterative Mixed Attention
Yutao Cui, Cheng Jiang, Limin Wang +1
Tracking often uses a multi-stage pipeline of feature extraction, target information integration, and bounding box estimation. To simplify this pipeline and unify the process of fe…