most citedSparse R-CNN: End-to-End Object Detection with Learnable Proposals

103 citations · 185 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV202236 cited

Learning Object-Language Alignments for Open-Vocabulary Object Detection

Chuang Lin, Peize Sun, Yi Jiang +5

Existing object detection methods are bounded in a fixed-set vocabulary by costly labeled data. When dealing with novel categories, the model has to be retrained with more bounding…

cs.CV20229 cited

Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders

Haosen Yang, Deng Huang, Bin Wen +5

Masked autoencoders (MAEs) have emerged recently as art self-supervised spatiotemporal representation learners. Inheriting from the image counterparts, however, existing video MAEs…

cs.CV202210 cited

Rethinking Resolution in the Context of Efficient Video Recognition

Chuofan Ma, Qiushan Guo, Yi Jiang +3

In this paper, we empirically study how to make the most of low-resolution frames for efficient video recognition. Existing methods mainly focus on developing compact networks or a…

cs.CV202227 cited

MetaFormer: A Unified Meta Framework for Fine-Grained Recognition

Qishuai Diao, Yi Jiang, Bin Wen +2

Fine-Grained Visual Classification(FGVC) is the task that requires recognizing the objects belonging to multiple subordinate categories of a super-category. Recent state-of-the-art…

cs.CV2020

TransTrack: Multiple Object Tracking with Transformer

Peize Sun, Jinkun Cao, Yi Jiang +5

In this work, we propose TransTrack, a simple but efficient scheme to solve the multiple object tracking problems. TransTrack leverages the transformer architecture, which is an at…

cs.CV2020

What Makes for End-to-End Object Detection?

Peize Sun, Yi Jiang, Enze Xie +4

Object detection has recently achieved a breakthrough for removing the last one non-differentiable component in the pipeline, Non-Maximum Suppression (NMS), and building up an end-…