Speed/accuracy trade-offs for modern convolutional object detectors
arXiv:1611.10012
Abstract
The goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detection systems. A number of successful systems have been proposed in recent years, but apples-to-apples comparisons are difficult due to different base feature extractors (e.g., VGG, Residual Networks), different default image resolutions, as well as different hardware and software platforms. We present a unified implementation of the Faster R-CNN [Ren et al., 2015], R-FCN [Dai et al., 2016] and SSD [Liu et al., 2015] systems, which we view as "meta-architectures" and trace out the speed/accuracy trade-off curve created by using alternative feature extractors and varying other critical parameters such as image size within each of these meta-architectures. On one extreme end of this spectrum where speed and memory are critical, we present a detector that achieves real time speeds and can be deployed on a mobile device. On the opposite end in which accuracy is critical, we present a detector that achieves state-of-the-art performance measured on the COCO detection task.
Accepted to CVPR 2017
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Caffe: Convolutional Architecture for Fast Feature Embedding
- DSSD : Deconvolutional Single Shot Detector
- An Analysis of Deep Neural Network Models for Practical Applications
- PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection
- Visual Discovery at Pinterest
Cited by in corpus (57)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Focal Loss for Dense Object Detection
- Deformable Convolutional Networks
- Cascade R-CNN: Delving into High Quality Object Detection
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Channel Pruning for Accelerating Very Deep Neural Networks
- Beyond Skip Connections: Top-Down Modulation for Object Detection
- Recognition of Ischaemia and Infection in Diabetic Foot Ulcers: Dataset and Techniques
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- An Implementation of Faster RCNN with Study for Region Sampling
- Receptive Field Block Net for Accurate and Fast Object Detection
- Single-Shot Refinement Neural Network for Object Detection
- Cascade R-CNN: High Quality Object Detection and Instance Segmentation
- Domain Adaptive Transfer Learning with Specialist Models
- Towards Accurate Multi-person Pose Estimation in the Wild
- Massive Exploration of Neural Machine Translation Architectures
- Relation Networks for Object Detection
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- A Survey on Deep Learning Methods for Robot Vision
- YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video
- YOLO-LITE: A Real-Time Object Detection Algorithm Optimized for Non-GPU Computers
- SFD: Single Shot Scale-invariant Face Detector
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- MegDet: A Large Mini-Batch Object Detector
- ChainerCV: a Library for Deep Learning in Computer Vision
- Deep-CEE I: Fishing for Galaxy Clusters with Deep Neural Nets
- The iNaturalist Species Classification and Detection Dataset
- Multi-Objective Automatic Machine Learning with AutoxgboostMC
- A Refined Deep Learning Architecture for Diabetic Foot Ulcers Detection
- Visual Discovery at Pinterest
- Okutama-Action: An Aerial View Video Dataset for Concurrent Human Action Detection
- Comparing Computing Platforms for Deep Learning on a Humanoid Robot
- A Framework of Transfer Learning in Object Detection for Embedded Systems
- Multi-scale Location-aware Kernel Representation for Object Detection
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report
- A Fog Robotics Approach to Deep Robot Learning: Application to Object Recognition and Grasp Planning in Surface Decluttering
- Relational Action Forecasting
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- Deep Regionlets for Object Detection
- Towards High Performance Video Object Detection
- Exploring Multi-Branch and High-Level Semantic Networks for Improving Pedestrian Detection
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- Fire SSD: Wide Fire Modules based Single Shot Detector on Edge Device
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- DeepLung: 3D Deep Convolutional Nets for Automated Pulmonary Nodule Detection and Classification
- Spatial Memory for Context Reasoning in Object Detection
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Sequence-guided protein structure determination using graph convolutional and recurrent networks
- A Novel Deep Neural Network Architecture for Mars Visual Navigation
- Weakly supervised one-stage vision and language disease detection using large scale pneumonia and pneumothorax studies
- Focal Loss Dense Detector for Vehicle Surveillance
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Working with scale: 2nd place solution to Product Detection in Densely Packed Scenes [Technical Report]
- Unifying data for fine-grained visual species classification
- Efficient Structured Pruning and Architecture Searching for Group Convolution
- A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection