259 citations · 385 across the 12 of their papers we have counts for
16 papers
detrex: Benchmarking Detection Transformers
Tianhe Ren, Shilong Liu, Feng Li +13
The DEtection TRansformer (DETR) algorithm has received considerable attention in the research community and is gradually emerging as a mainstream approach for object detection and…
A Strong and Reproducible Object Detector with Only Public Datasets
Tianhe Ren, Jianwei Yang, Shilong Liu +6
This work presents Focal-Stable-DINO, a strong and reproducible object detection model which achieves 64.6 AP on COCO val2017 and 64.8 AP on COCO test-dev using only 700M parameter…
Detection Transformer with Stable Matching
Shilong Liu, Tianhe Ren, Jiayu Chen +8
This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is cau…
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren +9
In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary obj…
CLIP4Caption ++: Multi-CLIP for Video Caption
Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng +2
This report describes our solution to the VALUE Challenge 2021 in the captioning task. Our solution, named CLIP4Caption++, is built on X-Linear/X-Transformer, which is an advanced…
Multi-modal Representation Learning for Video Advertisement Content Structuring
Daya Guo, Zhaoyang Zeng
Video advertisement content structuring aims to segment a given video advertisement and label each segment on various dimensions, such as presentation form, scene, and style. Diffe…