activity
20192023
most citedGrounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

259 citations · 385 across the 12 of their papers we have counts for

collaborators

16 papers

cs.CV2023★ 15 cited

detrex: Benchmarking Detection Transformers

Tianhe Ren, Shilong Liu, Feng Li +13

The DEtection TRansformer (DETR) algorithm has received considerable attention in the research community and is gradually emerging as a mainstream approach for object detection and…

cs.CV2023★ 4 cited

A Strong and Reproducible Object Detector with Only Public Datasets

Tianhe Ren, Jianwei Yang, Shilong Liu +6

This work presents Focal-Stable-DINO, a strong and reproducible object detection model which achieves 64.6 AP on COCO val2017 and 64.8 AP on COCO test-dev using only 700M parameter…

cs.CV2023★ 5 cited

Detection Transformer with Stable Matching

Shilong Liu, Tianhe Ren, Jiayu Chen +8

This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is cau…

cs.CV2023★ 259 cited

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Shilong Liu, Zhaoyang Zeng, Tianhe Ren +9

In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary obj…

cs.CV2021★ 4 cited

CLIP4Caption ++: Multi-CLIP for Video Caption

Mingkang Tang, Zhanyu Wang, Zhaoyang Zeng +2

This report describes our solution to the VALUE Challenge 2021 in the captioning task. Our solution, named CLIP4Caption++, is built on X-Linear/X-Transformer, which is an advanced…

cs.CV2021★ 3 cited

Multi-modal Representation Learning for Video Advertisement Content Structuring

Daya Guo, Zhaoyang Zeng

Video advertisement content structuring aims to segment a given video advertisement and label each segment on various dimensions, such as presentation form, scene, and style. Diffe…