Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
arXiv:2303.05499
Abstract
In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection is introducing language to a closed-set detector for open-set concept generalization. To effectively fuse language and vision modalities, we conceptually divide a closed-set detector into three phases and propose a tight fusion solution, which includes a feature enhancer, a language-guided query selection, and a cross-modality decoder for cross-modality fusion. While previous works mainly evaluate open-set object detection on novel categories, we propose to also perform evaluations on referring expression comprehension for objects specified with attributes. Grounding DINO performs remarkably well on all three settings, including benchmarks on COCO, LVIS, ODinW, and RefCOCO/+/g. Grounding DINO achieves a AP on the COCO detection zero-shot transfer benchmark, i.e., without any training data from COCO. It sets a new record on the ODinW zero-shot benchmark with a mean AP. Code will be available at \url{https://github.com/IDEA-Research/GroundingDINO}.
Code will be available at https://github.com/IDEA-Research/GroundingDINO
Cited by in corpus (11)
- BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- Language Models as Zero-Shot Trajectory Generators
- A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
- Sim-Suction: Learning a Suction Grasp Policy for Cluttered Environments Using a Synthetic Benchmark
- Open Scene Graphs for Open-World Object-Goal Navigation
- FM-Fusion: Instance-aware Semantic Mapping Boosted by Vision-Language Foundation Models
- Zero-shot detection of buildings in mobile LiDAR using Language Vision Model
- FashionFail: Addressing Failure Cases in Fashion Object Detection and Segmentation
- Utilizing Grounded SAM for self-supervised frugal camouflaged human detection
- Collaborating Foundation Models for Domain Generalized Semantic Segmentation