YOLO9000: Better, Faster, Stronger
arXiv:1612.08242
Abstract
We introduce YOLO9000, a state-of-the-art, real-time object detection system that can detect over 9000 object categories. First we propose various improvements to the YOLO detection method, both novel and drawn from prior work. The improved model, YOLOv2, is state-of-the-art on standard detection tasks like PASCAL VOC and COCO. At 67 FPS, YOLOv2 gets 76.8 mAP on VOC 2007. At 40 FPS, YOLOv2 gets 78.6 mAP, outperforming state-of-the-art methods like Faster RCNN with ResNet and SSD while still running significantly faster. Finally we propose a method to jointly train on object detection and classification. Using this method we train YOLO9000 simultaneously on the COCO detection dataset and the ImageNet classification dataset. Our joint training allows YOLO9000 to predict detections for object classes that don't have labelled detection data. We validate our approach on the ImageNet detection task. YOLO9000 gets 19.7 mAP on the ImageNet detection validation set despite only having detection data for 44 of the 200 classes. On the 156 classes not in COCO, YOLO9000 gets 16.0 mAP. But YOLO can detect more than just 200 classes; it predicts detections for more than 9000 different object categories. And it still runs in real-time.
References in corpus (2)
Cited by in corpus (74)
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- Single-Shot Refinement Neural Network for Object Detection
- Adversarial Examples that Fool Detectors
- Focus: Querying Large Video Datasets with Low Latency and Low Cost
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Region Proposal by Guided Anchoring
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- RefinedMPL: Refined Monocular PseudoLiDAR for 3D Object Detection in Autonomous Driving
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- Scaling Video Analytics on Constrained Edge Nodes
- 3D-A-Nets: 3D Deep Dense Descriptor for Volumetric Shapes with Adversarial Networks
- Residual Features and Unified Prediction Network for Single Stage Detection
- Robust and High Performance Face Detector
- Lightweight Convolutional Neural Network with Gaussian-based Grasping Representation for Robotic Grasping Detection
- Highrisk Prediction from Electronic Medical Records via Deep Attention Networks
- Self-supervised Moving Vehicle Tracking with Stereo Sound
- Extend the shallow part of Single Shot MultiBox Detector via Convolutional Neural Network
- PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
- Knowledge Projection for Deep Neural Networks
- Feature Intertwiner for Object Detection
- Efficient Golf Ball Detection and Tracking Based on Convolutional Neural Networks and Kalman Filter
- Improving Multiple Object Tracking with Optical Flow and Edge Preprocessing
- Building Robust Deep Neural Networks for Road Sign Detection
- UAV Visual Teach and Repeat Using Only Semantic Object Features
- Weaving Multi-scale Context for Single Shot Detector
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Towards High Performance Video Object Detection
- 360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
- Grab: Fast and Accurate Sensor Processing for Cashier-Free Shopping
- Joint Face Detection and Facial Motion Retargeting for Multiple Faces
- Building Proactive Voice Assistants: When and How (not) to Interact
- GestARLite: An On-Device Pointing Finger Based Gestural Interface for Smartphones and Video See-Through Head-Mounts
- Learning with Hierarchical Complement Objective
- Matrix and tensor decompositions for training binary neural networks
- Spot the Difference by Object Detection
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Developing a Compressed Object Detection Model based on YOLOv4 for Deployment on Embedded GPU Platform of Autonomous System
- Automated flow for compressing convolution neural networks for efficient edge-computation with FPGA
- Using Deep Networks for Drone Detection
- TextTubes for Detecting Curved Text in the Wild
- Spot Evasion Attacks: Adversarial Examples for License Plate Recognition Systems with Convolutional Neural Networks
- Symbol, Conversational, and Societal Grounding with a Toy Robot
- Single Shot Multitask Pedestrian Detection and Behavior Prediction
- Simultaneous x, y Pixel Estimation and Feature Extraction for Multiple Small Objects in a Scene: A Description of the ALIEN Network
- Progressive Representation Adaptation for Weakly Supervised Object Localization
- Plug & Play Convolutional Regression Tracker for Video Object Detection
- LIP: Learning Instance Propagation for Video Object Segmentation
- Semantic Segmentation and Object Detection Towards Instance Segmentation: Breast Tumor Identification
- Phasic dopamine release identification using ensemble of AlexNet
- AirPen: A Touchless Fingertip Based Gestural Interface for Smartphones and Head-Mounted Devices
- Mask R-CNN Based Object Detection for Intelligent Wireless Power Transfer
- The Role of Compute in Autonomous Aerial Vehicles
- Exploring Effectiveness of Inter-Microtask Qualification Tests in Crowdsourcing
- DeepSIC: Deep Semantic Image Compression
- Evaluating robustness of You Only Hear Once(YOHO) Algorithm on noisy audios in the VOICe Dataset
- Serious Games Application for Memory Training Using Egocentric Images
- Convolutional Neural Networks for Real-Time Localization and Classification in Feedback Digital Microscopy
- Gaussian Filter in CRF Based Semantic Segmentation
- Scalable Object Detection for Stylized Objects
- Latent Cognizance: What Machine Really Learns
- Clique: Spatiotemporal Object Re-identification at the City Scale
- Proactive Network Maintenance using Fast, Accurate Anomaly Localization and Classification on 1-D Data Series
- Place-specific Background Modeling Using Recursive Autoencoders
- MyFood: A Food Segmentation and Classification System to Aid Nutritional Monitoring
- Learning Pixel Representations for Generic Segmentation
- A Distributed Framework to Orchestrate Video Analytics Applications
- Adversarial Semantic Contour for Object Detection
- Learning to Parse Wireframes in Images of Man-Made Environments
- Learning to Predict the 3D Layout of a Scene
- LBGP: Learning Based Goal Planning for Autonomous Following in Front
- Efficient Pipelines for Vision-Based Context Sensing