Feature Pyramid Networks for Object Detection
arXiv:1612.03144
Abstract
Feature pyramids are a basic component in recognition systems for detecting objects at different scales. But recent deep learning object detectors have avoided pyramid representations, in part because they are compute and memory intensive. In this paper, we exploit the inherent multi-scale, pyramidal hierarchy of deep convolutional networks to construct feature pyramids with marginal extra cost. A top-down architecture with lateral connections is developed for building high-level semantic feature maps at all scales. This architecture, called a Feature Pyramid Network (FPN), shows significant improvement as a generic feature extractor in several applications. Using FPN in a basic Faster R-CNN system, our method achieves state-of-the-art single-model results on the COCO detection benchmark without bells and whistles, surpassing all existing single-model entries including those from the COCO 2016 challenge winners. In addition, our method can run at 5 FPS on a GPU and thus is a practical and accurate solution to multi-scale object detection. Code will be made publicly available.
Cited by in corpus (173)
- YOLOv3: An Incremental Improvement
- Focal Loss for Dense Object Detection
- Automatic Ship Detection of Remote Sensing Images from Google Earth in Complex Scenes Based on Multi-Scale Rotation Dense Feature Pyramid Networks
- Deformable Convolutional Networks
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- A Review of Object Detection Models based on Convolutional Neural Network
- Deep Learning Enables Automatic Detection and Segmentation of Brain Metastases on Multi-Sequence MRI
- Slimmable Neural Networks
- Semantic Instance Segmentation via Deep Metric Learning
- Learning a Rotation Invariant Detector with Rotatable Bounding Box
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- Differentiable Learning-to-Normalize via Switchable Normalization
- CornerNet: Detecting Objects as Paired Keypoints
- Face Attention Network: An Effective Face Detector for the Occluded Faces
- Toward Transformer-Based Object Detection
- Adapting Mask-RCNN for Automatic Nucleus Segmentation
- Single-Shot Refinement Neural Network for Object Detection
- Rethinking ImageNet Pre-training
- Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection
- BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
- Orthographic Feature Transform for Monocular 3D Object Detection
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection
- Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment
- Detecting and Recognizing Human-Object Interactions
- ExFuse: Enhancing Feature Fusion for Semantic Segmentation
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- Square Kilometre Array Science Data Challenge 1: analysis and results
- A Survey on Deep Learning Methods for Robot Vision
- WordSup: Exploiting Word Annotations for Character based Text Detection
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- Towards High Performance Video Object Detection for Mobiles
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- Where are the Masks: Instance Segmentation with Image-level Supervision
- Tiling and Stitching Segmentation Output for Remote Sensing: Basic Challenges and Recommendations
- Learning to Compose Dynamic Tree Structures for Visual Contexts
- Adversarial Sparse-View CBCT Artifact Reduction
- Using Deep Learning for Segmentation and Counting within Microscopy Data
- Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking
- Learning Markov Clustering Networks for Scene Text Detection
- MegDet: A Large Mini-Batch Object Detector
- Class-incremental Learning via Deep Model Consolidation
- Iterative Visual Reasoning Beyond Convolutions
- Text Detection and Recognition in the Wild: A Review
- Effect of Annotation Errors on Drone Detection with YOLOv3
- Feature Agglomeration Networks for Single Stage Face Detection
- Cell Tracking via Proposal Generation and Selection
- Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships
- Residual Features and Unified Prediction Network for Single Stage Detection
- Temporal Pyramid Network for Action Recognition
- Crowd counting via scale-adaptive convolutional neural network
- Multi-hierarchical Independent Correlation Filters for Visual Tracking
- SIXray : A Large-scale Security Inspection X-ray Benchmark for Prohibited Item Discovery in Overlapping Images
- Repulsion Loss: Detecting Pedestrians in a Crowd
- PaveSAM Segment Anything for Pavement Distress
- A survey of Object Classification and Detection based on 2D/3D data
- ZeroQ: A Novel Zero Shot Quantization Framework
- Sketch-R2CNN: An Attentive Network for Vector Sketch Recognition
- TextMountain: Accurate Scene Text Detection via Instance Segmentation
- Learning Discriminative Motion Features Through Detection
- Seeing Small Faces from Robust Anchor's Perspective
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Accurate Monocular Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving
- A Mask-RCNN Baseline for Probabilistic Object Detection
- Fusing Bird View LIDAR Point Cloud and Front View Camera Image for Deep Object Detection
- Contextual Multi-Scale Region Convolutional 3D Network for Activity Detection
- Learning to Segment Every Thing
- PolyTransform: Deep Polygon Transformer for Instance Segmentation
- The Fishnet Open Images Database: A Dataset for Fish Detection and Fine-Grained Categorization in Fisheries
- Multi-scale Location-aware Kernel Representation for Object Detection
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- Distilling Knowledge via Knowledge Review
- Optimizing Video Object Detection via a Scale-Time Lattice
- DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- Blind Predicting Similar Quality Map for Image Quality Assessment
- Domain Adaptation from Synthesis to Reality in Single-model Detector for Video Smoke Detection
- EyeNet: A Multi-Task Network for Off-Axis Eye Gaze Estimation and User Understanding
- The Devil is in the Decoder: Classification, Regression and GANs
- MultiResolution Attention Extractor for Small Object Detection
- NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection
- FOTS: Fast Oriented Text Spotting with a Unified Network
- Deep Concept-wise Temporal Convolutional Networks for Action Localization
- Weaving Multi-scale Context for Single Shot Detector
- Efficient Palm-Line Segmentation with U-Net Context Fusion Module
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection
- ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel Matching
- PointIT: A Fast Tracking Framework Based on 3D Instance Segmentation
- Deep Regionlets for Object Detection
- Can we cover navigational perception needs of the visually impaired by panoptic segmentation?
- Towards large-scale, automated, accurate detection of CCTV camera objects using computer vision. Applications and implications for privacy, safety, and cybersecurity. (Preprint)
- Fine-Grained Visual Classification of Plant Species In The Wild: Object Detection as A Reinforced Means of Attention
- Recurrent Scale Approximation for Object Detection in CNN
- Towards High Performance Video Object Detection
- Context-Aware Single-Shot Detector
- 3D Object Detection From LiDAR Data Using Distance Dependent Feature Extraction
- Detection and Attention: Diagnosing Pulmonary Lung Cancer from CT by Imitating Physicians
- Semi-supervised Learning: Fusion of Self-supervised, Supervised Learning, and Multimodal Cues for Tactical Driver Behavior Detection
- Two-stream Convolutional Networks for Multi-frame Face Anti-spoofing
- Zero-Annotation Object Detection with Web Knowledge Transfer
- Improvement of Classification in One-Stage Detector
- Attention Mechanisms for Object Recognition with Event-Based Cameras
- HySTER: A Hybrid Spatio-Temporal Event Reasoner
- Image Recognition Using Scale Recurrent Neural Networks
- Pixel-Attentive Policy Gradient for Multi-Fingered Grasping in Cluttered Scenes
- Beyond Trade-off: Accelerate FCN-based Face Detector with Higher Accuracy
- Adversarial Occlusion-aware Face Detection
- Density Map Guided Object Detection in Aerial Images
- Crowd Scene Analysis by Output Encoding
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- Video Representation Learning and Latent Concept Mining for Large-scale Multi-label Video Classification
- AIO-P: Expanding Neural Performance Predictors Beyond Image Classification
- Frustum VoxNet for 3D object detection from RGB-D or Depth images
- Multi Scale Supervised 3D U-Net for Kidney and Tumor Segmentation
- Pedestrian Detection with Autoregressive Network Phases
- RePr: Improved Training of Convolutional Filters
- Deep Feature Pyramid Reconfiguration for Object Detection
- Dynamic Filtering with Large Sampling Field for ConvNets
- Detecting Heads using Feature Refine Net and Cascaded Multi-Scale Architecture
- Spatial Memory for Context Reasoning in Object Detection
- Object Detection with Mask-based Feature Encoding
- End-to-End Entity Detection with Proposer and Regressor
- Subspace Match Probably Does Not Accurately Assess the Similarity of Learned Representations
- Toward Robotic Weed Control: Detection of Nutsedge Weed in Bermudagrass Turf Using Inaccurate and Insufficient Training Data
- Image-based Virtual Fitting Room
- DIPN: Deep Interaction Prediction Network with Application to Clutter Removal
- Deep Imbalanced Attribute Classification using Visual Attention Aggregation
- Key Frame Proposal Network for Efficient Pose Estimation in Videos
- Investigation of a Machine learning methodology for the SKA pulsar search pipeline
- Using Computer Vision to Automate Hand Detection and Tracking of Surgeon Movements in Videos of Open Surgery
- A Locating Model for Pulmonary Tuberculosis Diagnosis in Radiographs
- 2nd Place Solution in Google AI Open Images Object Detection Track 2019
- Instance Scale Normalization for image understanding
- TextTubes for Detecting Curved Text in the Wild
- Multi Target Tracking by Learning from Generalized Graph Differences
- Cascading Neural Network Methodology for Artificial Intelligence-Assisted Radiographic Detection and Classification of Lead-Less Implanted Electronic Devices within the Chest
- Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
- PØDA: Prompt-driven Zero-shot Domain Adaptation
- PSDet: Efficient and Universal Parking Slot Detection
- Towards Understanding the Effectiveness of Attention Mechanism
- Transformer-F: A Transformer network with effective methods for learning universal sentence representation
- Semantic Image Cropping
- FA-RPN: Floating Region Proposals for Face Detection
- Fully convolutional Siamese neural networks for buildings damage assessment from satellite images
- Feature Selective Networks for Object Detection
- TLGAN: document Text Localization using Generative Adversarial Nets
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- Focal Loss Dense Detector for Vehicle Surveillance
- KPNet: Towards Minimal Face Detector
- Object Detection on Single Monocular Images through Canonical Correlation Analysis
- Recognition of Russian traffic signs in winter conditions. Solutions of the "Ice Vision" competition winners
- A Multi-Task Learning Approach for Meal Assessment
- Auto-calibration Method Using Stop Signs for Urban Autonomous Driving Applications
- 2nd Place Solution to Instance Segmentation of IJCAI 3D AI Challenge 2020
- A Global to Local Double Embedding Method for Multi-person Pose Estimation
- Convolutions for Spatial Interaction Modeling
- Non-local RoIs for Instance Segmentation
- Learning to Predict the 3D Layout of a Scene
- Detector-in-Detector: Multi-Level Analysis for Human-Parts
- Rapidly Adapting Moment Estimation
- Dense Multiscale Feature Fusion Pyramid Networks for Object Detection in UAV-Captured Images
- Convolutional Recurrent Network for Road Boundary Extraction
- DAGMapper: Learning to Map by Discovering Lane Topology
- Hierarchical Recurrent Attention Networks for Structured Online Maps
- Identity Enhanced Residual Image Denoising
- A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection
- DETCID: Detection of Elongated Touching Cells with Inhomogeneous Illumination using a Deep Adversarial Network
- A Light-Weight Object Detection Framework with FPA Module for Optical Remote Sensing Imagery
- Multi-Scale Gradual Integration CNN for False Positive Reduction in Pulmonary Nodule Detection
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Safe Augmentation: Learning Task-Specific Transformations from Data