441 citations · 2.5k across the 54 of their papers we have counts for
12 papers · 1 filter
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Yi Wang, Kunchang Li, Yizhuo Li +14
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on…
VLG: General Video Recognition with Web Textual Knowledge
Jintao Lin, Zhaoyang Liu, Wenhai Wang +2
Video recognition in an open and dynamic world is quite challenging, as we need to handle different settings such as close-set, long-tail, few-shot and open-set. By leveraging sema…
UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Kunchang Li, Yali Wang, Yinan He +4
Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term v…
Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding
Fengyuan Shi, Ruopeng Gao, Weilin Huang +1
Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) s…
BasicTAD: an Astounding RGB-Only Baseline for Temporal Action Detection
Min Yang, Guo Chen, Yin-Dong Zheng +2
Temporal action detection (TAD) is extensively studied in the video understanding community by generally following the object detection pipeline in images. However, complex designs…
AdaMixer: A Fast-Converging Query-Based Object Detector
Ziteng Gao, Limin Wang, Bing Han +1
Traditional object detectors employ the dense paradigm of scanning over locations and scales in an image. The recent query-based object detectors break this convention by decoding…