272 citations · 534 across the 14 of their papers we have counts for
25 papers
VLG: General Video Recognition with Web Textual Knowledge
Jintao Lin, Zhaoyang Liu, Wenhai Wang +2
Video recognition in an open and dynamic world is quite challenging, as we need to handle different settings such as close-set, long-tail, few-shot and open-set. By leveraging sema…
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
Hao Li, Jinguo Zhu, Xiaohu Jiang +8
Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to elimi…
On Efficient Reinforcement Learning for Full-length Game of StarCraft II
Ruo-Ze Liu, Zhen-Jia Pang, Zhou-Yu Meng +3
StarCraft II (SC2) poses a grand challenge for reinforcement learning (RL), of which the main difficulties include huge state space, varying action space, and a long time horizon.…
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality
Xiang Li, Wenhai Wang, Lingfeng Yang +1
Masked AutoEncoder (MAE) has recently led the trends of visual self-supervision area by an elegant asymmetric encoder-decoder design, which significantly optimizes both the pre-tra…
WegFormer: Transformers for Weakly Supervised Semantic Segmentation
Chunmeng Liu, Enze Xie, Wenjia Wang +3
Although convolutional neural networks (CNNs) have achieved remarkable progress in weakly supervised semantic segmentation (WSSS), the effective receptive field of CNN is insuffici…
ARTS: Eliminating Inconsistency between Text Detection and Recognition with Auto-Rectification Text Spotter
Humen Zhong, Jun Tang, Wenhai Wang +3
Recent approaches for end-to-end text spotting have achieved promising results. However, most of the current spotters were plagued by the inconsistency problem between text detecti…