Beyond Skip Connections: Top-Down Modulation for Object Detection
arXiv:1612.06851
Abstract
In recent years, we have seen tremendous progress in the field of object detection. Most of the recent improvements have been achieved by targeting deeper feedforward networks. However, many hard object categories such as bottle, remote, etc. require representation of fine details and not just coarse, semantic representations. But most of these fine details are lost in the early convolutional layers. What we need is a way to incorporate finer details from lower layers into the detection architecture. Skip connections have been proposed to combine high-level and low-level features, but we argue that selecting the right features from low-level requires top-down contextual information. Inspired by the human visual pathway, in this paper we propose top-down modulations as a way to incorporate fine details into the detection framework. Our approach supplements the standard bottom-up, feedforward ConvNet with a top-down modulation (TDM) network, connected using lateral connections. These connections are responsible for the modulation of lower layer filters, and the top-down network handles the selection and integration of contextual information and low-level features. The proposed TDM architecture provides a significant boost on the COCO testdev benchmark, achieving 28.6 AP for VGG16, 35.2 AP for ResNet101, and 37.3 for InceptionResNetv2 network, without any bells and whistles (e.g., multi-scale, iterative box refinement, etc.).
References in corpus (4)
Cited by in corpus (76)
- YOLOv3: An Incremental Improvement
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Focal Loss for Dense Object Detection
- A Survey of Deep Learning-based Object Detection
- FCOS: Fully Convolutional One-Stage Object Detection
- Object Detection in 20 Years: A Survey
- Path Aggregation Network for Instance Segmentation
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Deep Learning for Generic Object Detection: A Survey
- Pyramid Attention Network for Semantic Segmentation
- CenterNet: Keypoint Triplets for Object Detection
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation
- CornerNet: Detecting Objects as Paired Keypoints
- SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- Backbones-Review: Feature Extraction Networks for Deep Learning and Deep Reinforcement Learning Approaches
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Single-Shot Refinement Neural Network for Object Detection
- Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection
- Review of data analysis in vision inspection of power lines with an in-depth discussion of deep learning technology
- Scale-Aware Trident Networks for Object Detection
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- A Survey on Deep Learning Methods for Robot Vision
- Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
- FCOS: A simple and strong anchor-free object detector
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- A Deep Learning Approach for Pose Estimation from Volumetric OCT Data
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- RefineDetLite: A Lightweight One-stage Object Detection Framework for CPU-only Devices
- Iterative Visual Reasoning Beyond Convolutions
- Recent Advances in Deep Learning for Object Detection
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
- Intrinsic Relationship Reasoning for Small Object Detection
- Consistent Optimization for Single-Shot Object Detection
- Target Driven Instance Detection
- Actor-Centric Relation Network
- Feature Intertwiner for Object Detection
- Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific Localization
- An Analysis of Scale Invariance in Object Detection - SNIP
- Occluded Prohibited Items Detection: an X-ray Security Inspection Benchmark and De-occlusion Attention Module
- Multi-scale Location-aware Kernel Representation for Object Detection
- DPNet: Dynamic Pooling Network for Tiny Object Detection
- Guided Attention Network for Object Detection and Counting on Drones
- MFPN: A Novel Mixture Feature Pyramid Network of Multiple Architectures for Object Detection
- Single Pixel Reconstruction for One-stage Instance Segmentation
- MTL-NAS: Task-Agnostic Neural Architecture Search towards General-Purpose Multi-Task Learning
- IoU-uniform R-CNN: Breaking Through the Limitations of RPN
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- ScratchDet: Training Single-Shot Object Detectors from Scratch
- LapNet : Automatic Balanced Loss and Optimal Assignment for Real-Time Dense Object Detection
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- PointAtrousGraph: Deep Hierarchical Encoder-Decoder with Point Atrous Convolution for Unorganized 3D Points
- ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel Matching
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- iffDetector: Inference-aware Feature Filtering for Object Detection
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- Concatenated Feature Pyramid Network for Instance Segmentation
- Outline Objects using Deep Reinforcement Learning
- Spatial Memory for Context Reasoning in Object Detection
- Occlusion-shared and Feature-separated Network for Occlusion Relationship Reasoning
- Deep Feature Pyramid Reconfiguration for Object Detection
- Fast Efficient Object Detection Using Selective Attention
- Multiple Anchor Learning for Visual Object Detection
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- Spatial Priming for Detecting Human-Object Interactions
- HR-RCNN: Hierarchical Relational Reasoning for Object Detection
- Focal Loss Dense Detector for Vehicle Surveillance
- Dive Deeper Into Box for Object Detection
- Multi-Grid Redundant Bounding Box Annotation for Accurate Object Detection
- Trident Pyramid Networks: The importance of processing at the feature pyramid level for better object detection
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Feature Flow: In-network Feature Flow Estimation for Video Object Detection
- ScarfNet: Multi-scale Features with Deeply Fused and Redistributed Semantics for Enhanced Object Detection
- POD: Practical Object Detection with Scale-Sensitive Network
- RMOPP: Robust Multi-Objective Post-Processing for Effective Object Detection