DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
arXiv:2203.03605
Abstract
We present DINO (\textbf{D}ETR with \textbf{I}mproved de\textbf{N}oising anch\textbf{O}r boxes), a state-of-the-art end-to-end object detector. % in this paper. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a mixed query selection method for anchor initialization, and a look forward twice scheme for box prediction. DINO achieves AP in epochs and AP in epochs on COCO with a ResNet-50 backbone and multi-scale features, yielding a significant improvement of \textbf{AP} and \textbf{AP}, respectively, compared to DN-DETR, the previous best DETR-like model. DINO scales well in both model size and data size. Without bells and whistles, after pre-training on the Objects365 dataset with a SwinL backbone, DINO obtains the best results on both COCO \texttt{val2017} (\textbf{AP}) and \texttt{test-dev} (\textbf{AP}). Compared to other models on the leaderboard, DINO significantly reduces its model size and pre-training data size while achieving better results. Our code will be available at \url{https://github.com/IDEACVR/DINO}.
Cited by in corpus (27)
- UnitModule: A Lightweight Joint Image Enhancement Module for Underwater Object Detection
- Physical Adversarial Attack meets Computer Vision: A Decade Survey
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- A Multi-task Framework for Infrared Small Target Detection and Segmentation
- Oriented object detection in optical remote sensing images using deep learning: a survey
- Autonomous Advanced Aerial Mobility -- An End-to-end Autonomy Framework for UAVs and Beyond
- E-UAV: An Edge-based Energy-Efficient Object Detection System for Unmanned Aerial Vehicles
- SpecDETR: A transformer-based hyperspectral point object detection network
- Vehicle Perception from Satellite
- Deep Learning-Based Connector Detection for Robotized Assembly of Automotive Wire Harnesses
- A Multimodal Dataset and Benchmark for Radio Galaxy and Infrared Host Detection
- MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction
- Foundation Model-Based Apple Ripeness and Size Estimation for Selective Harvesting
- Focused Decoding Enables 3D Anatomical Detection by Transformers
- Taming Detection Transformers for Medical Object Detection
- Visual inspection for illicit items in X-ray images using Deep Learning
- Multiple Object Tracking in Video SAR: A Benchmark and Tracking Baseline
- VME: A Satellite Imagery Dataset and Benchmark for Detecting Vehicles in the Middle East and Beyond
- Parametric Primitive Analysis of CAD Sketches with Vision Transformer
- ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents
- STRIDE: Street View-based Environmental Feature Detection and Pedestrian Collision Prediction
- Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
- OAH-Net: A Deep Neural Network for Hologram Reconstruction of Off-axis Digital Holographic Microscope
- Illicit object detection in X-ray images using Vision Transformers
- Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
- Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- Wild Berry image dataset collected in Finnish forests and peatlands using drones