Deformable DETR: Deformable Transformers for End-to-End Object Detection
arXiv:2010.04159
Abstract
DETR has been recently proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance. However, it suffers from slow convergence and limited feature spatial resolution, due to the limitation of Transformer attention modules in processing image feature maps. To mitigate these issues, we proposed Deformable DETR, whose attention modules only attend to a small set of key sampling points around a reference. Deformable DETR can achieve better performance than DETR (especially on small objects) with 10 times less training epochs. Extensive experiments on the COCO benchmark demonstrate the effectiveness of our approach. Code is released at https://github.com/fundamentalvision/Deformable-DETR.
ICLR 2021 Oral
References in corpus (12)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Longformer: The Long-Document Transformer
- Linformer: Self-Attention with Linear Complexity
- Generating Long Sequences with Sparse Transformers
- Axial Attention in Multidimensional Transformers
- Reformer: The Efficient Transformer
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Big Bird: Transformers for Longer Sequences
- Image Transformer
- Sparse Sinkhorn Attention
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
Cited by in corpus (166)
- Transformers in Vision: A Survey
- TransTrack: Multiple Object Tracking with Transformer
- Centralized Feature Pyramid for Object Detection
- A Comprehensive Survey on Deep Graph Representation Learning
- Waste detection in Pomerania: non-profit project for detecting waste in environment
- End-to-end Temporal Action Detection with Transformer
- LocalViT: Analyzing Locality in Vision Transformers
- TransCrowd: weakly-supervised crowd counting with transformers
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- Radar-Camera Fusion for Object Detection and Semantic Segmentation in Autonomous Driving: A Comprehensive Review
- Building extraction with vision transformer
- DepthFormer: Exploiting Long-Range Correlation and Local Information for Accurate Monocular Depth Estimation
- Efficient Training of Audio Transformers with Patchout
- UNETR: Transformers for 3D Medical Image Segmentation
- TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Restormer: Efficient Transformer for High-Resolution Image Restoration
- CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation
- Deep Learning Approaches on Image Captioning: A Review
- CBNet: A Composite Backbone Network Architecture for Object Detection
- CvT: Introducing Convolutions to Vision Transformers
- Short and Long Range Relation Based Spatio-Temporal Transformer for Micro-Expression Recognition
- Probabilistic two-stage detection
- Vision Transformer with Attentive Pooling for Robust Facial Expression Recognition
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- ResT: An Efficient Transformer for Visual Recognition
- Alpha-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression
- Toward Transformer-Based Object Detection
- Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point Supervision
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- Pre-Trained Image Processing Transformer
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- End-to-End Object Detection with Adaptive Clustering Transformer
- DPT: Deformable Patch-based Transformer for Visual Recognition
- Transformer with Peak Suppression and Knowledge Guidance for Fine-grained Image Recognition
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Adjacent-Level Feature Cross-Fusion With 3-D CNN for Remote Sensing Image Change Detection
- CycleMLP: A MLP-like Architecture for Dense Prediction
- Efficient Self-supervised Vision Transformers for Representation Learning
- Part-guided Relational Transformers for Fine-grained Visual Recognition
- TrTr: Visual Tracking with Transformer
- Rethinking Spatial Dimensions of Vision Transformers
- ConTNet: Why not use convolution and transformer at the same time?
- CMT: Convolutional Neural Networks Meet Vision Transformers
- TBFormer: Two-Branch Transformer for Image Forgery Localization
- CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
- SOLQ: Segmenting Objects by Learning Queries
- ISTR: End-to-End Instance Segmentation with Transformers
- Towards Holistic Surgical Scene Understanding
- Vision Transformer Pruning
- SODFormer: Streaming Object Detection with Transformer Using Events and Frames
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Deep Blind Super-Resolution for Satellite Video
- Transformer-based Map Matching Model with Limited Ground-Truth Data using Transfer-Learning Approach
- Relation Matters: Foreground-aware Graph-based Relational Reasoning for Domain Adaptive Object Detection
- PPT Fusion: Pyramid Patch Transformerfor a Case Study in Image Fusion
- A Comprehensive Review of Modern Object Segmentation Approaches
- Video Joint Modelling Based on Hierarchical Transformer for Co-summarization
- Vision Transformers with Patch Diversification
- Refiner: Refining Self-attention for Vision Transformers
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement
- QueryProp: Object Query Propagation for High-Performance Video Object Detection
- DETR for Crowd Pedestrian Detection
- Few-shot Action Recognition with Prototype-centered Attentive Learning
- TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
- Dynamic Label Assignment for Object Detection by Combining Predicted IoUs and Anchor IoUs
- RQFormer: Rotated Query Transformer for End-to-End Oriented Object Detection
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
- Few-Shot Segmentation via Cycle-Consistent Transformer
- MATIS: Masked-Attention Transformers for Surgical Instrument Segmentation
- Segmenting Transparent Object in the Wild with Transformer
- Deep Learning Technique for Human Parsing: A Survey and Outlook
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- Video Crowd Localization with Multi-focus Gaussian Neighborhood Attention and a Large-Scale Benchmark
- Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
- PSViT: Better Vision Transformer via Token Pooling and Attention Sharing
- TokenPose: Learning Keypoint Tokens for Human Pose Estimation
- Automatic Rail Component Detection Based on AttnConv-Net
- Fully Transformer-Equipped Architecture for End-to-End Referring Video Object Segmentation
- HODOR: High-level Object Descriptors for Object Re-segmentation in Video Learned from Static Images
- OH-Former: Omni-Relational High-Order Transformer for Person Re-Identification
- Graph-Segmenter: Graph Transformer with Boundary-aware Attention for Semantic Segmentation
- A Multimodal Dataset and Benchmark for Radio Galaxy and Infrared Host Detection
- Towards Unsupervised Domain Adaptation via Domain-Transformer
- Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation
- View-Disentangled Transformer for Brain Lesion Detection
- Diverse Image Inpainting with Bidirectional and Autoregressive Transformers
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction
- Multi-Compound Transformer for Accurate Biomedical Image Segmentation
- GiT: Graph Interactive Transformer for Vehicle Re-identification
- PU-Transformer: Point Cloud Upsampling Transformer
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- DeepArUco++: Improved detection of square fiducial markers in challenging lighting conditions
- Sound Event Detection Transformer: An Event-based End-to-End Model for Sound Event Detection
- Understanding the computational demands underlying visual reasoning
- Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
- Table Detection for Visually Rich Document Images
- Multi-Modal Learning for AU Detection Based on Multi-Head Fused Transformers
- Improving 3D Object Detection with Channel-wise Transformer
- OffRoadTranSeg: Semi-Supervised Segmentation using Transformers on OffRoad environments
- Detecting Gender Bias in Transformer-based Models: A Case Study on BERT
- Softmax-free Linear Transformers
- HiFT: Hierarchical Feature Transformer for Aerial Tracking
- HIH: Towards More Accurate Face Alignment via Heatmap in Heatmap
- Parameter-Efficient Fine-Tuning of Large Pretrained Models for Instance Segmentation Tasks
- Spatially Consistent Representation Learning
- Deep neural networks approach to microbial colony detection -- a comparative analysis
- Distance-Aware Occlusion Detection with Focused Attention
- End-to-End Trainable Multi-Instance Pose Estimation with Transformers
- TransMed: Transformers Advance Multi-modal Medical Image Classification
- Multitask Learning in Minimally Invasive Surgical Vision: A Review
- 3D Object Detection with Pointformer
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- Lightning Fast Video Anomaly Detection via Adversarial Knowledge Distillation
- A Lightweight NMS-free Framework for Real-time Visual Fault Detection System of Freight Trains
- Towards Biologically Plausible Convolutional Networks
- Transformer Meets Convolution: A Bilateral Awareness Network for Semantic Segmentation of Very Fine Resolution Urban Scene Images
- Cycle Self-Training for Semi-Supervised Object Detection with Distribution Consistency Reweighting
- CAT: Cross-Attention Transformer for One-Shot Object Detection
- Vision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification
- CrowdSim2: an Open Synthetic Benchmark for Object Detectors
- DBIA: Data-free Backdoor Injection Attack against Transformer Networks
- Parametric Primitive Analysis of CAD Sketches with Vision Transformer
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting
- Deep-learning Assisted Detection and Quantification of (oo)cysts of Giardia and Cryptosporidium on Smartphone Microscopy Images
- You Only Look Bottom-Up for Monocular 3D Object Detection
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Transformer-based Network for RGB-D Saliency Detection
- TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
- Towards High-Quality Temporal Action Detection with Sparse Proposals
- End-to-End Chess Recognition
- Split and Connect: A Universal Tracklet Booster for Multi-Object Tracking
- UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
- LaRa: Latents and Rays for Multi-Camera Bird's-Eye-View Semantic Segmentation
- End-to-End Dense Video Grounding via Parallel Regression
- Towards Few-Annotation Learning for Object Detection: Are Transformer-based Models More Efficient ?
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- Post-Training Quantization for Vision Transformer
- Structured Sparse R-CNN for Direct Scene Graph Generation
- Sampling Equivariant Self-attention Networks for Object Detection in Aerial Images
- Rethinking Open-Set Object Detection: Issues, a New Formulation, and Taxonomy
- Transformer for Polyp Detection
- Dense Object Detection Based on De-homogenized Queries
- Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text
- Spectral Unsupervised Domain Adaptation for Visual Recognition
- I2C2W: Image-to-Character-to-Word Transformers for Accurate Scene Text Recognition
- Efficient Video Transformers with Spatial-Temporal Token Selection
- The MIS Check-Dam Dataset for Object Detection and Instance Segmentation Tasks
- SparseDet: Towards End-to-End 3D Object Detection
- Guiding Query Position and Performing Similar Attention for Transformer-Based Detection Heads
- Attend to Who You Are: Supervising Self-Attention for Keypoint Detection and Instance-Aware Association
- Region-Adaptive Deformable Network for Image Quality Assessment
- CT-block: a novel local and global features extractor for point cloud
- OadTR: Online Action Detection with Transformers
- MaskBEV: Joint Object Detection and Footprint Completion for Bird's-eye View 3D Point Clouds
- Transformer-based Dual Relation Graph for Multi-label Image Recognition
- Trident Pyramid Networks: The importance of processing at the feature pyramid level for better object detection
- A Lightweight Graph Transformer Network for Human Mesh Reconstruction from 2D Human Pose
- Learning Dynamic Compact Memory Embedding for Deformable Visual Object Tracking
- A Comparison of Deep Learning Methods for Cell Detection in Digital Cytology
- Synthesizing Photorealistic Images with Deep Generative Learning