Scalable, High-Quality Object Detection
arXiv:1412.1441
Abstract
Current high-quality object detection approaches use the scheme of salience-based object proposal methods followed by post-classification using deep convolutional features. This spurred recent research in improving object proposal methods. However, domain agnostic proposal generation has the principal drawback that the proposals come unranked or with very weak ranking, making it hard to trade-off quality for running time. This raises the more fundamental question of whether high-quality proposal generation requires careful engineering or can be derived just from data alone. We demonstrate that learning-based proposal methods can effectively match the performance of hand-engineered methods while allowing for very efficient runtime-quality trade-offs. Using the multi-scale convolutional MultiBox (MSC-MultiBox) approach, we substantially advance the state-of-the-art on the ILSVRC 2014 detection challenge data set, with mAP for a single model and mAP for an ensemble of two models. MSC-Multibox significantly improves the proposal quality over its predecessor MultiBox~method: AP increases from to for the ILSVRC detection challenge. Finally, we demonstrate improved bounding-box recall compared to Multiscale Combinatorial Grouping with less proposals on the Microsoft-COCO data set.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Learning to Segment Object Candidates
- Multiscale Combinatorial Grouping for Image Segmentation and Object Proposal Generation
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- How good are detection proposals, really?
Cited by in corpus (56)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- ParseNet: Looking Wider to See Better
- What makes for effective detection proposals?
- Learning to Segment Object Candidates
- Object Detection in 20 Years: A Survey
- Deformable Convolutional Networks
- An Empirical Evaluation of Deep Learning on Highway Driving
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Recent Advances in Convolutional Neural Networks
- Pyramid Scene Parsing Network
- Deep Learning for Generic Object Detection: A Survey
- Speed/accuracy trade-offs for modern convolutional object detectors
- A Light CNN for Deep Face Representation with Noisy Labels
- Image Super-Resolution Using Deep Convolutional Networks
- Attention for Fine-Grained Categorization
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Object Detection with Deep Learning: A Review
- DenseCap: Fully Convolutional Localization Networks for Dense Captioning
- Rule Extraction Algorithm for Deep Neural Networks: A Review
- Learning to Refine Object Segments
- Data Distillation: Towards Omni-Supervised Learning
- End-to-end Learning of Action Detection from Frame Glimpses in Videos
- Automatic Spatially-aware Fashion Concept Discovery
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Deep learning in radiology: an overview of the concepts and a survey of the state of the art
- DeepBox: Learning Objectness with Convolutional Networks
- Crafting GBD-Net for Object Detection
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- Boosting Convolutional Features for Robust Object Proposals
- Exploring Person Context and Local Scene Context for Object Detection
- Efficient Object Detection for High Resolution Images
- Anchor Box Optimization for Object Detection
- End-to-End Object Detection with Fully Convolutional Network
- AttentionNet: Aggregating Weak Directions for Accurate Object Detection
- Learning Region Features for Object Detection
- Learning Robust Deep Face Representation
- Large Scale Business Discovery from Street Level Imagery
- Learning Fine-grained Features via a CNN Tree for Large-scale Classification
- Unsupervised Learning of Edges
- Active Object Localization in Visual Situations
- Learning to detect and localize many objects from few examples
- Learning to Learn Relation for Important People Detection in Still Images
- SeGAN: Segmenting and Generating the Invisible
- Towards Adversarially Robust Object Detection
- Damage GAN: A Generative Model for Imbalanced Data
- LiDAR and Camera Detection Fusion in a Real Time Industrial Multi-Sensor Collision Avoidance System
- Distill-2MD-MTL: Data Distillation based on Multi-Dataset Multi-Domain Multi-Task Frame Work to Solve Face Related Tasksks, Multi Task Learning, Semi-Supervised Learning
- What leads to generalization of object proposals?
- Learnable Histogram: Statistical Context Features for Deep Neural Networks
- Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction Recognition
- Semantic Segmentation via Highly Fused Convolutional Network with Multiple Soft Cost Functions
- Fusion of an Ensemble of Augmented Image Detectors for Robust Object Detection
- Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object Detection
- Self-supervised Transfer Learning for Instance Segmentation through Physical Interaction
- Improved Super-Resolution Convolution Neural Network for Large Images
- Diversity in Object Proposals