activity
20192022
most citedDecoupled Spatial-Temporal Transformer for Video Inpainting

47 citations · 76 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV20223 cited

Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information

Weijie Su, Xizhou Zhu, Chenxin Tao +7

To effectively exploit the potential of large-scale models, various pre-training strategies supported by massive data from different sources are proposed, including supervised pre-…

cs.CV20226 cited

BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

Chenyu Yang, Yuntao Chen, Hao Tian +9

We present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and better suits modern image backbones. Existing state-of-the-art BEV detect…

cs.CV20217 cited

FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting

Rui Liu, Hanming Deng, Yangyi Huang +6

Transformer, as a strong and flexible architecture for modelling long-range relations, has been widely explored in vision tasks. However, when used in video inpainting that require…

cs.CV202147 cited

Decoupled Spatial-Temporal Transformer for Video Inpainting

Rui Liu, Hanming Deng, Yangyi Huang +6

Video inpainting aims to fill the given spatiotemporal holes with realistic appearance but is still a challenging task even with prosperous deep learning approaches. Recent works i…

cs.CV2020

Deformable DETR: Deformable Transformers for End-to-End Object Detection

Xizhou Zhu, Weijie Su, Lewei Lu +3

DETR has been recently proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance. However, it suffers from slow conv…

cs.CV202013 cited

1st Place Solution of LVIS Challenge 2020: A Good Box is not a Guarantee of a Good Mask

Jingru Tan, Gang Zhang, Hanming Deng +4

This article introduces the solutions of the team lvisTraveler for LVIS Challenge 2020. In this work, two characteristics of LVIS dataset are mainly considered: the long-tailed dis…