BoxCars: Improving Fine-Grained Recognition of Vehicles using 3-D Bounding Boxes in Traffic Surveillance
arXiv:1703.00686 · doi:10.1109/TITS.2018.2799228
Abstract
In this paper, we focus on fine-grained recognition of vehicles mainly in traffic surveillance applications. We propose an approach that is orthogonal to recent advancements in fine-grained recognition (automatic part discovery and bilinear pooling). In addition, in contrast to other methods focused on fine-grained recognition of vehicles, we do not limit ourselves to a frontal/rear viewpoint, but allow the vehicles to be seen from any viewpoint. Our approach is based on 3-D bounding boxes built around the vehicles. The bounding box can be automatically constructed from traffic surveillance data. For scenarios where it is not possible to use precise construction, we propose a method for an estimation of the 3-D bounding box. The 3-D bounding box is used to normalize the image viewpoint by "unpacking" the image into a plane. We also propose to randomly alter the color of the image and add a rectangle with random noise to a random position in the image during the training of convolutional neural networks (CNNs). We have collected a large fine-grained vehicle data set BoxCars116k, with 116k images of vehicles from various viewpoints taken by numerous surveillance cameras. We performed a number of experiments, which show that our proposed method significantly improves CNN classification accuracy (the accuracy is increased by up to 12% points and the error is reduced by up to 50% compared with CNNs without the proposed modifications). We also show that our method outperforms the state-of-the-art methods for fine-grained recognition.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Using Deep Learning and Google Street View to Estimate the Demographic Makeup of the US
- Evolving Boxes for Fast Vehicle Detection
- Object Contour Detection with a Fully Convolutional Encoder-Decoder Network
Cited by in corpus (8)
- Detection of 3D Bounding Boxes of Vehicles Using Perspective Transformation for Accurate Speed Measurement
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Efficient Vision-based Vehicle Speed Estimation
- A Fine-Grained Vehicle Detection (FGVD) Dataset for Unconstrained Roads
- Improving Vehicle Re-Identification using CNN Latent Spaces: Metrics Comparison and Track-to-track Extension
- Robust, Extensible, and Fast: Teamed Classifiers for Vehicle Tracking and Vehicle Re-ID in Multi-Camera Networks
- Traffic Camera Calibration via Vehicle Vanishing Point Detection
- Automated Object Behavioral Feature Extraction for Potential Risk Analysis based on Video Sensor