DenseBox: Unifying Landmark Localization with End to End Object Detection
arXiv:1509.04874
Abstract
How can a single fully convolutional neural network (FCN) perform on object detection? We introduce DenseBox, a unified end-to-end FCN framework that directly predicts bounding boxes and object class confidences through all locations and scales of an image. Our contribution is two-fold. First, we show that a single FCN, if designed and optimized carefully, can detect multiple different objects extremely accurately and efficiently. Second, we show that when incorporating with landmark localization during multi-task learning, DenseBox further improves object detection accuray. We present experimental results on public benchmark datasets including MALF face detection and KITTI car detection, that indicate our DenseBox is the state-of-the-art system for detecting challenging objects such as faces and cars.
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Going Deeper with Convolutions
- ParseNet: Looking Wider to See Better
- Learning to Segment Object Candidates
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- Heterogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network
- What is Holding Back Convnets for Detection?
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
Cited by in corpus (100)
- UnitBox: An Advanced Object Detection Network
- FCOS: Fully Convolutional One-Stage Object Detection
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- RepPoints: Point Set Representation for Object Detection
- Face Attention Network: An Effective Face Detector for the Occluded Faces
- One-Stage Cascade Refinement Networks for Infrared Small Target Detection
- EAST: An Efficient and Accurate Scene Text Detector
- Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- Probabilistic and Geometric Depth: Detecting Objects in Perspective
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Object Detection with Deep Learning: A Review
- A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task Learning
- Orthographic Feature Transform for Monocular 3D Object Detection
- DirectPose: Direct End-to-End Multi-Person Pose Estimation
- RepPoints V2: Verification Meets Regression for Object Detection
- Combining Data-driven and Model-driven Methods for Robust Facial Landmark Detection
- CityPersons: A Diverse Dataset for Pedestrian Detection
- Face Detection Using Improved Faster RCNN
- Revisiting Feature Alignment for One-stage Object Detection
- PolarMask: Single Shot Instance Segmentation with Polar Representation
- Region Proposal by Guided Anchoring
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines
- Acquisition of Localization Confidence for Accurate Object Detection
- PointFusion: Deep Sensor Fusion for 3D Bounding Box Estimation
- PixelNet: Towards a General Pixel-level Architecture
- Deep Direct Regression for Multi-Oriented Scene Text Detection
- AIParsing: Anchor-free Instance-level Human Parsing
- PDNet: Toward Better One-Stage Object Detection With Prediction Decoupling
- WordSup: Exploiting Word Annotations for Character based Text Detection
- Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation
- FCOS: A simple and strong anchor-free object detector
- PyramidBox++: High Performance Detector for Finding Tiny Face
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- VarifocalNet: An IoU-aware Dense Object Detector
- TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes
- CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection
- PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection
- Scale-Aware Face Detection
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection
- iQIYI-VID: A Large Dataset for Multi-modal Person Identification
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
- Repulsion Loss: Detecting Pedestrians in a Crowd
- Towards Interpretable and Robust Hand Detection via Pixel-wise Prediction
- SFace: An Efficient Network for Face Detection in Large Scale Variations
- Siamese Box Adaptive Network for Visual Tracking
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Adapted Center and Scale Prediction: More Stable and More Accurate
- An Anchor-Free Region Proposal Network for Faster R-CNN based Text Detection Approaches
- BorderDet: Border Feature for Dense Object Detection
- Dense Regression Network for Video Grounding
- IoU-aware Single-stage Object Detector for Accurate Localization
- Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
- Slender Object Detection: Diagnoses and Improvements
- MSR: Multi-Scale Shape Regression for Scene Text Detection
- Soft Anchor-Point Object Detection
- Training-Time-Friendly Network for Real-Time Object Detection
- Mask R-CNN with Pyramid Attention Network for Scene Text Detection
- Object Detection in Video with Spatial-temporal Context Aggregation
- PolarMask++: Enhanced Polar Representation for Single-Shot Instance Segmentation and Beyond
- Training Bit Fully Convolutional Network for Fast Semantic Segmentation
- Category-wise Attack: Transferable Adversarial Examples for Anchor Free Object Detection
- Multi-object Tracking via End-to-end Tracklet Searching and Ranking
- CenterMask: single shot instance segmentation with point representation
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- OTA: Optimal Transport Assignment for Object Detection
- PIXOR: Real-time 3D Object Detection from Point Clouds
- DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
- DuBox: No-Prior Box Objection Detection via Residual Dual Scale Detectors
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Beyond Trade-off: Accelerate FCN-based Face Detector with Higher Accuracy
- SaccadeNet: A Fast and Accurate Object Detector
- ReLaText: Exploiting Visual Relationships for Arbitrary-Shaped Scene Text Detection with Graph Convolutional Networks
- Weakly-Supervised Arbitrary-Shaped Text Detection with Expectation-Maximization Algorithm
- RRPN++: Guidance Towards More Accurate Scene Text Detection
- Foreground-Background Imbalance Problem in Deep Object Detectors: A Review
- Never Mind the Bounding Boxes, Here's the SAND Filters
- TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network
- SADet: Learning An Efficient and Accurate Pedestrian Detector
- Resisting Crowd Occlusion and Hard Negatives for Pedestrian Detection in the Wild
- Temporal Self-Ensembling Teacher for Semi-Supervised Object Detection
- Disentangle Your Dense Object Detector
- Detecting Text in the Wild with Deep Character Embedding Network
- High Performance Visual Object Tracking with Unified Convolutional Networks
- Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object Detectors
- Multi-Grid Redundant Bounding Box Annotation for Accurate Object Detection
- Dive Deeper Into Box for Object Detection
- Fast Local Attack: Generating Local Adversarial Examples for Object Detectors
- Deep Point-wise Prediction for Action Temporal Proposal
- Accurate Scene Text Detection through Border Semantics Awareness and Bootstrapping
- A Light-Weight Object Detection Framework with FPA Module for Optical Remote Sensing Imagery
- Modulating Localization and Classification for Harmonized Object Detection
- Pixel-Semantic Revise of Position Learning A One-Stage Object Detector with A Shared Encoder-Decoder
- Objects as Extreme Points
- EOLO: Embedded Object Segmentation only Look Once
- Boundary-Aware Dense Feature Indicator for Single-Stage 3D Object Detection from Point Clouds
- AGSFCOS: Based on attention mechanism and Scale-Equalizing pyramid network of object detection