Deformable Convolutional Networks
arXiv:1703.06211
Abstract
Convolutional neural networks (CNNs) are inherently limited to model geometric transformations due to the fixed geometric structures in its building modules. In this work, we introduce two new modules to enhance the transformation modeling capacity of CNNs, namely, deformable convolution and deformable RoI pooling. Both are based on the idea of augmenting the spatial sampling locations in the modules with additional offsets and learning the offsets from target tasks, without additional supervision. The new modules can readily replace their plain counterparts in existing CNNs and can be easily trained end-to-end by standard back-propagation, giving rise to deformable convolutional networks. Extensive experiments validate the effectiveness of our approach on sophisticated vision tasks of object detection and semantic segmentation. The code would be released.
References in corpus (8)
- Feature Pyramid Networks for Object Detection
- Simultaneous Detection and Segmentation
- Dilated Residual Networks
- DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection
- Deformable Part Models are Convolutional Neural Networks
- Harmonic Networks: Deep Translation and Rotation Equivariance
- Invariant Scattering Convolution Networks
- Inverse Compositional Spatial Transformer Networks
Cited by in corpus (179)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Attention Mechanisms in Computer Vision: A Survey
- Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
- Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
- Gabor Convolutional Networks
- Cascade R-CNN: Delving into High Quality Object Detection
- FSSD: Feature Fusion Single Shot Multibox Detector
- Structure-sensitive Multi-scale Deep Neural Network for Low-Dose CT Denoising
- Deep Learning for Generic Object Detection: A Survey
- Generative Image Inpainting with Contextual Attention
- Deformable ConvNets v2: More Deformable, Better Results
- R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object
- CornerNet: Detecting Objects as Paired Keypoints
- SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
- Multi-Scale Supervised 3D U-Net for Kidneys and Kidney Tumor Segmentation
- Deep Learning in Mobile and Wireless Networking: A Survey
- Receptive Field Block Net for Accurate and Fast Object Detection
- Simple Baselines for Human Pose Estimation and Tracking
- Recent Advances in Deep Learning: An Overview
- Single-Shot Refinement Neural Network for Object Detection
- Crowd Counting by Adaptively Fusing Predictions from an Image Pyramid
- Flow-Guided Feature Aggregation for Video Object Detection
- What Do We Understand About Convolutional Networks?
- Selective Kernel Networks
- Object Detection with Deep Learning: A Review
- Stacked Deconvolutional Network for Semantic Segmentation
- TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
- Learning to Measure Change: Fully Convolutional Siamese Metric Networks for Scene Change Detection
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- Revisiting Feature Alignment for One-stage Object Detection
- EDVR: Video Restoration with Enhanced Deformable Convolutional Networks
- Video Instance Segmentation using Inter-Frame Communication Transformers
- Learning Human-Object Interactions by Graph Parsing Neural Networks
- Soft Sampling for Robust Object Detection
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Efficient Video Object Segmentation via Network Modulation
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- ST-GAN: Spatial Transformer Generative Adversarial Networks for Image Compositing
- Learning Deformable Kernels for Image and Video Denoising
- Scene Text Recognition from Two-Dimensional Perspective
- IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- Towards High Performance Video Object Detection for Mobiles
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
- Impression Network for Video Object Detection
- WIDER Face and Pedestrian Challenge 2018: Methods and Results
- FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling Network
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- MegDet: A Large Mini-Batch Object Detector
- The Herbarium Challenge 2019 Dataset
- ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations
- Compressing 3DCNNs Based on Tensor Train Decomposition
- Progressive Sparse Local Attention for Video object detection
- Constructing Fast Network through Deconstruction of Convolution
- Rotation Equivariance and Invariance in Convolutional Neural Networks
- Image Segmentation and Classification for Sickle Cell Disease using Deformable U-Net
- Image Inpainting for Irregular Holes Using Partial Convolutions
- Dynamic Instance Normalization for Arbitrary Style Transfer
- Learning Discriminative Motion Features Through Detection
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Seeing Small Faces from Robust Anchor's Perspective
- Contrastive Learning for Compact Single Image Dehazing
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- Dense Transformer Networks
- Graph-Based Global Reasoning Networks
- Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Integral Human Pose Regression
- HRDNet: High-resolution Detection Network for Small Objects
- Towards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks
- M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- Indices Matter: Learning to Index for Deep Image Matting
- Geometry meets semantics for semi-supervised monocular depth estimation
- Control Distance IoU and Control Distance IoU Loss Function for Better Bounding Box Regression
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- Temporal Interlacing Network
- YH Technologies at ActivityNet Challenge 2018
- Learning Region Features for Object Detection
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Compact Global Descriptor for Neural Networks
- Geometric robustness of deep networks: analysis and improvement
- NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
- Super-Resolution with Deep Adaptive Image Resampling
- Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
- ScratchDet: Training Single-Shot Object Detectors from Scratch
- BoundarySqueeze: Image Segmentation as Boundary Squeezing
- Trimmed Action Recognition, Dense-Captioning Events in Videos, and Spatio-temporal Action Localization with Focus on ActivityNet Challenge 2019
- Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks
- Comparison-Based Convolutional Neural Networks for Cervical Cell/Clumps Detection in the Limited Data Scenario
- IPG-Net: Image Pyramid Guidance Network for Small Object Detection
- Deep Regionlets for Object Detection
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- 1st Place Solution for ICDAR 2021 Competition on Mathematical Formula Detection
- Learning from Noisy Anchors for One-stage Object Detection
- 2nd Place Solution for Waymo Open Dataset Challenge -- 2D Object Detection
- SOGNet: Scene Overlap Graph Network for Panoptic Segmentation
- Towards High Performance Video Object Detection
- Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
- ArbiText: Arbitrary-Oriented Text Detection in Unconstrained Scene
- DeRPN: Taking a further step toward more general object detection
- Irregular Convolutional Neural Networks
- SAN: Learning Relationship between Convolutional Features for Multi-Scale Object Detection
- Beyond Domain Adaptation: Unseen Domain Encapsulation via Universal Non-volume Preserving Models
- Grounded Objects and Interactions for Video Captioning
- A spatiotemporal model with visual attention for video classification
- Learning a Unified Sample Weighting Network for Object Detection
- Large-Scale Object Detection in the Wild from Imbalanced Multi-Labels
- Deformable Deep Convolutional Generative Adversarial Network in Microwave Based Hand Gesture Recognition System
- Outline Objects using Deep Reinforcement Learning
- Object Detection based on Region Decomposition and Assembly
- FlatteNet: A Simple Versatile Framework for Dense Pixelwise Prediction
- Neuron-level Selective Context Aggregation for Scene Segmentation
- Improving the Resolution of CNN Feature Maps Efficiently with Multisampling
- Propose-and-Attend Single Shot Detector
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Dynamic Filtering with Large Sampling Field for ConvNets
- Inability of spatial transformations of CNN feature maps to support invariant recognition
- Learning to Caricature via Semantic Shape Transform
- Isometric Transformation Invariant Graph-based Deep Neural Network
- Instance Scale Normalization for image understanding
- Inter-BMV: Interpolation with Block Motion Vectors for Fast Semantic Segmentation on Video
- STAS: Adaptive Selecting Spatio-Temporal Deep Features for Improving Bias Correction on Precipitation
- Geometric Operator Convolutional Neural Network
- Augmenting Proposals by the Detector Itself
- Decoupled IoU Regression for Object Detection
- Affine Self Convolution
- Spatially-Adaptive Filter Units for Deep Neural Networks
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement
- Multiple receptive fields and small-object-focusing weakly-supervised segmentation network for fast object detection
- Local Binary Pattern Networks
- Object Detection with Mask-based Feature Encoding
- I3DOL: Incremental 3D Object Learning without Catastrophic Forgetting
- FA-RPN: Floating Region Proposals for Face Detection
- Low Pass Filter for Anti-aliasing in Temporal Action Localization
- Generating Unrestricted Adversarial Examples via Three Parameters
- Feature Selective Networks for Object Detection
- Mass Displacement Networks
- Statistical transformer networks: learning shape and appearance models via self supervision
- ShuffleDet: Real-Time Vehicle Detection Network in On-board Embedded UAV Imagery
- Doppler-Radar Based Hand Gesture Recognition System Using Convolutional Neural Networks
- An Efficient Accelerator Design Methodology for Deformable Convolutional Networks
- Deformable Part Networks
- Video Panoptic Segmentation
- A New Ensemble Adversarial Attack Powered by Long-term Gradient Memories
- FineNet: Frame Interpolation and Enhancement for Face Video Deblurring
- Cross-View Image Synthesis with Deformable Convolution and Attention Mechanism
- PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
- G-RCN: Optimizing the Gap between Classification and Localization Tasks for Object Detection
- Correlation Propagation Networks for Scene Text Detection
- Fine-grained Image-to-Image Transformation towards Visual Recognition
- Semantic Image Retrieval by Uniting Deep Neural Networks and Cognitive Architectures
- Plug & Play Convolutional Regression Tracker for Video Object Detection
- Focal Loss Dense Detector for Vehicle Surveillance
- ContourRender: Detecting Arbitrary Contour Shape For Instance Segmentation In One Pass
- Know Your Surroundings: Panoramic Multi-Object Tracking by Multimodality Collaboration
- 1st Place Solutions for UG2+ Challenge 2021 -- (Semi-)supervised Face detection in the low light condition
- Cyclic orthogonal convolutions for long-range integration of features
- A Top-down Approach to Articulated Human Pose Estimation and Tracking
- Large-scale mammography CAD with Deformable Conv-Nets
- Learning Equivariant Representations
- Affinity Derivation and Graph Merge for Instance Segmentation
- Towards Fine-grained Large Object Segmentation 1st Place Solution to 3D AI Challenge 2020 -- Instance Segmentation Track
- 2nd Place Solution to Instance Segmentation of IJCAI 3D AI Challenge 2020
- Ada-Segment: Automated Multi-loss Adaptation for Panoptic Segmentation
- DeepKey: Towards End-to-End Physical Key Replication From a Single Photograph
- Orientation Convolutional Networks for Image Recognition
- BAN: Focusing on Boundary Context for Object Detection
- 3rd Place Solution for Short-video Face Parsing Challenge
- Non-local RoI for Cross-Object Perception
- Deformable Stacked Structure for Named Entity Recognition
- GAIA: A Transfer Learning System of Object Detection that Fits Your Needs