Visual Compiler: Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator
arXiv:1612.05234
Abstract
We introduce the concept of a Visual Compiler that generates a scene specific pedestrian detector and pose estimator without any pedestrian observations. Given a single image and auxiliary scene information in the form of camera parameters and geometric layout of the scene, the Visual Compiler first infers geometrically and photometrically accurate images of humans in that scene through the use of computer graphics rendering. Using these renders we learn a scene-and-region specific spatially-varying fully convolutional neural network, for simultaneous detection, pose estimation and segmentation of pedestrians. We demonstrate that when real human annotated data is scarce or non-existent, our data generation strategy can provide an excellent solution for bootstrapping human detection and pose estimation. Experimental results show that our approach outperforms off-the-shelf state-of-the-art pedestrian detectors and pose estimators that are trained on real data.
submitted to CVPR 2017
References in corpus (8)
- DeepPose: Human Pose Estimation via Deep Neural Networks
- Fully Convolutional Networks for Semantic Segmentation
- Stacked Hourglass Networks for Human Pose Estimation
- Convolutional Pose Machines
- Identity Mappings in Deep Residual Networks
- Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views
- Learning Complexity-Aware Cascades for Deep Pedestrian Detection
- Filtered Channel Features for Pedestrian Detection
Cited by in corpus (5)
- 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics
- GNN3DMOT: Graph Neural Network for 3D Multi-Object Tracking with Multi-Feature Learning
- When We First Met: Visual-Inertial Person Localization for Co-Robot Rendezvous
- AutoSelect: Automatic and Dynamic Detection Selection for 3D Multi-Object Tracking
- GroundNet: Monocular Ground Plane Normal Estimation with Geometric Consistency