Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
arXiv:1512.04143
Abstract
It is well known that contextual and multi-scale representations are important for accurate visual recognition. In this paper we present the Inside-Outside Net (ION), an object detector that exploits information both inside and outside the region of interest. Contextual information outside the region of interest is integrated using spatial recurrent neural networks. Inside, we use skip pooling to extract information at multiple scales and levels of abstraction. Through extensive experiments we evaluate the design space and provide readers with an overview of what tricks of the trade are important. ION improves state-of-the-art on PASCAL VOC 2012 object detection from 73.9% to 76.4% mAP. On the new and more challenging MS COCO dataset, we improve state-of-art-the from 19.7% to 33.1% mAP. In the 2015 MS COCO Detection Challenge, our ION model won the Best Student Entry and finished 3rd place overall. As intuition suggests, our detection results provide strong evidence that context and multi-scale representations improve small object detection.
References in corpus (13)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- ParseNet: Looking Wider to See Better
- Visualizing and Understanding Recurrent Networks
- Fully Convolutional Networks for Semantic Segmentation
- What makes for effective detection proposals?
- Learning to Segment Object Candidates
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks
- Analyzing the Performance of Multilayer Neural Networks for Object Recognition
- Pedestrian Detection with Unsupervised Multi-Stage Feature Learning
Cited by in corpus (36)
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- YOLO9000: Better, Faster, Stronger
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Speed/accuracy trade-offs for modern convolutional object detectors
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- Perceptual Generative Adversarial Networks for Small Object Detection
- Learning Feature Pyramids for Human Pose Estimation
- PixelNet: Towards a General Pixel-level Architecture
- Modeling Context in Referring Expressions
- Combining the Best of Convolutional Layers and Recurrent Layers: A Hybrid Network for Semantic Segmentation
- A deep architecture for unified aesthetic prediction
- MegDet: A Large Mini-Batch Object Detector
- Face Detection using Deep Learning: An Improved Faster RCNN Approach
- Feature Agglomeration Networks for Single Stage Face Detection
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- Supervised Transformer Network for Efficient Face Detection
- Adaptive Object Detection Using Adjacency and Zoom Prediction
- Single-Shot Object Detection with Enriched Semantics
- Deconvolutional Feature Stacking for Weakly-Supervised Semantic Segmentation
- An Analysis of Scale Invariance in Object Detection - SNIP
- Multi-scale Location-aware Kernel Representation for Object Detection
- The Role of Context Selection in Object Detection
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- Weaving Multi-scale Context for Single Shot Detector
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- Multi-stage Object Detection with Group Recursive Learning
- Learning to detect and localize many objects from few examples
- Spatial Memory for Context Reasoning in Object Detection
- Object-Level Context Modeling For Scene Classification with Context-CNN
- Not Using the Car to See the Sidewalk: Quantifying and Controlling the Effects of Context in Classification and Segmentation
- Deep neural networks can be improved using human-derived contextual expectations
- Precise Box Score: Extract More Information from Datasets to Improve the Performance of Face Detection
- Deep Markov Random Field for Image Modeling
- Fast Learning and Prediction for Object Detection using Whitened CNN Features
- Feature Selective Networks for Object Detection
- Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural Networks