R-FCN: Object Detection via Region-based Fully Convolutional Networks
arXiv:1605.06409
Abstract
We present region-based, fully convolutional networks for accurate and efficient object detection. In contrast to previous region-based detectors such as Fast/Faster R-CNN that apply a costly per-region subnetwork hundreds of times, our region-based detector is fully convolutional with almost all computation shared on the entire image. To achieve this goal, we propose position-sensitive score maps to address a dilemma between translation-invariance in image classification and translation-variance in object detection. Our method can thus naturally adopt fully convolutional image classifier backbones, such as the latest Residual Networks (ResNets), for object detection. We show competitive results on the PASCAL VOC datasets (e.g., 83.6% mAP on the 2007 set) with the 101-layer ResNet. Meanwhile, our result is achieved at a test-time speed of 170ms per image, 2.5-20x faster than the Faster R-CNN counterpart. Code is made publicly available at: https://github.com/daijifeng001/r-fcn
Tech report
Cited by in corpus (558)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- A Survey of Deep Learning Techniques for Autonomous Driving
- Focal Loss for Dense Object Detection
- Recent advances and clinical applications of deep learning in medical image analysis
- Lunar impact craters identification and age estimation with Chang'E data by deep and transfer learning
- Gliding vertex on the horizontal bounding box for multi-oriented object detection
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Richer Convolutional Features for Edge Detection
- Deep Learning in Video Multi-Object Tracking: A Survey
- Machine learning in acoustics: theory and applications
- Automatic Ship Detection of Remote Sensing Images from Google Earth in Complex Scenes Based on Multi-Scale Rotation Dense Feature Pyramid Networks
- Object Detection in 20 Years: A Survey
- YOLACT++: Better Real-time Instance Segmentation
- YOLOP: You Only Look Once for Panoptic Driving Perception
- Deep Continuous Fusion for Multi-Sensor 3D Object Detection
- Deformable Convolutional Networks
- Cascade R-CNN: Delving into High Quality Object Detection
- FSSD: Feature Fusion Single Shot Multibox Detector
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- Detecting Mammals in UAV Images: Best Practices to address a substantially Imbalanced Dataset with Deep Learning
- A Simple Semi-Supervised Learning Framework for Object Detection
- The EuroCity Persons Dataset: A Novel Benchmark for Object Detection
- DetNet: A Backbone network for Object Detection
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler Divergence
- Deep Learning for Generic Object Detection: A Survey
- PP-YOLO: An Effective and Efficient Implementation of Object Detector
- Detecting Curve Text in the Wild: New Dataset and New Solution
- Rethinking Rotated Object Detection with Gaussian Wasserstein Distance Loss
- PVANET: Deep but Lightweight Neural Networks for Real-time Object Detection
- Mitigating Adversarial Effects Through Randomization
- Few-Example Object Detection with Model Communication
- Iterative Filter Adaptive Network for Single Image Defocus Deblurring
- Hybrid Task Cascade for Instance Segmentation
- Position Detection and Direction Prediction for Arbitrary-Oriented Ships via Multitask Rotation Region Convolutional Neural Network
- CenterNet: Keypoint Triplets for Object Detection
- TSM: Temporal Shift Module for Efficient Video Understanding
- Learning a Rotation Invariant Detector with Rotatable Bounding Box
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- DeeperLab: Single-Shot Image Parser
- CornerNet-Lite: Efficient Keypoint Based Object Detection
- From Handcrafted to Deep Features for Pedestrian Detection: A Survey
- Efficient DETR: Improving End-to-End Object Detector with Dense Prior
- R3Det: Refined Single-Stage Detector with Feature Refinement for Rotating Object
- Beyond Part Models: Person Retrieval with Refined Part Pooling (and a Strong Convolutional Baseline)
- DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic
- RepPoints: Point Set Representation for Object Detection
- SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
- Neonatal seizure detection from raw multi-channel EEG using a fully convolutional architecture
- DOTA: A Large-scale Dataset for Object Detection in Aerial Images
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Towards Accurate One-Stage Object Detection with AP-Loss
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Receptive Field Block Net for Accurate and Fast Object Detection
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- Rain rendering for evaluating and improving robustness to bad weather
- Simple Baselines for Human Pose Estimation and Tracking
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Single-Shot Refinement Neural Network for Object Detection
- Do Adversarially Robust ImageNet Models Transfer Better?
- Multivariate Confidence Calibration for Object Detection
- Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection
- Bottom-up Object Detection by Grouping Extreme and Center Points
- Flow-Guided Feature Aggregation for Video Object Detection
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Sparse R-CNN: End-to-End Object Detection with Learnable Proposals
- PolarDet: A Fast, More Precise Detector for Rotated Target in Aerial Images
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- Retina U-Net: Embarrassingly Simple Exploitation of Segmentation Supervision for Medical Object Detection
- Review of data analysis in vision inspection of power lines with an in-depth discussion of deep learning technology
- Adversarial Examples for Semantic Segmentation and Object Detection
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- Segmentation of cell-level anomalies in electroluminescence images of photovoltaic modules
- Finding beans in burgers: Deep semantic-visual embedding with localization
- AP-Loss for Accurate One-Stage Object Detection
- Scale-Aware Trident Networks for Object Detection
- Object Detection with Deep Learning: A Review
- xView: Objects in Context in Overhead Imagery
- SkyNet: a Hardware-Efficient Method for Object Detection and Tracking on Embedded Systems
- PVANet: Lightweight Deep Neural Networks for Real-time Object Detection
- Advances and Applications of Computer Vision Techniques in Vehicle Trajectory Generation and Surrogate Traffic Safety Indicators
- Learning Rich Features for Image Manipulation Detection
- ME R-CNN: Multi-Expert R-CNN for Object Detection
- IENet: Interacting Embranchment One Stage Anchor Free Detector for Orientation Aerial Object Detection
- Robust Adversarial Perturbation on Deep Proposal-based Models
- Towards Accurate Multi-person Pose Estimation in the Wild
- Attention-guided Context Feature Pyramid Network for Object Detection
- Gradient Harmonized Single-stage Detector
- Relation Networks for Object Detection
- RON: Reverse Connection with Objectness Prior Networks for Object Detection
- Face Detection Using Improved Faster RCNN
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- Dynamic Head: Unifying Object Detection Heads with Attentions
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- Learning Modulated Loss for Rotated Object Detection
- Volumetric Attention for 3D Medical Image Segmentation and Detection
- MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
- Real-time Plant Health Assessment Via Implementing Cloud-based Scalable Transfer Learning On AWS DeepLens
- Region Proposal by Guided Anchoring
- Vehicle Detection of Multi-source Remote Sensing Data Using Active Fine-tuning Network
- Privacy Protection in Street-View Panoramas using Depth and Multi-View Imagery
- Style Aggregated Network for Facial Landmark Detection
- BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation
- Background Learnable Cascade for Zero-Shot Object Detection
- PixelLink: Detecting Scene Text via Instance Segmentation
- Soft Sampling for Robust Object Detection
- DEFT: Detection Embeddings for Tracking
- Small-scale Pedestrian Detection Based on Somatic Topology Localization and Temporal Feature Aggregation
- Enhancing Object Detection for Autonomous Driving by Optimizing Anchor Generation and Addressing Class Imbalance
- Integrated Object Detection and Tracking with Tracklet-Conditioned Detection
- Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
- Global Aggregation then Local Distribution in Fully Convolutional Networks
- Accurate Single Stage Detector Using Recurrent Rolling Convolution
- Chart-Text: A Fully Automated Chart Image Descriptor
- Distilling Object Detectors with Task Adaptive Regularization
- Prime Sample Attention in Object Detection
- SolarNet: A Deep Learning Framework to Map Solar Power Plants In China From Satellite Imagery
- Learning deep structured active contours end-to-end
- ThunderNet: Towards Real-time Generic Object Detection
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
- Multi-Oriented Scene Text Detection via Corner Localization and Region Segmentation
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- MobileDets: Searching for Object Detection Architectures for Mobile Accelerators
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Rethinking Classification and Localization for Object Detection
- A Vision-based Social Distancing and Critical Density Detection System for COVID-19
- Sequence Level Semantics Aggregation for Video Object Detection
- AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
- Weakly Supervised Instance Segmentation using Class Peak Response
- RelationNet++: Bridging Visual Representations for Object Detection via Transformer Decoder
- IncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
- Fast object detection in compressed JPEG Images
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- Towards High Performance Video Object Detection for Mobiles
- Dynamic Label Assignment for Object Detection by Combining Predicted IoUs and Anchor IoUs
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- GFD-SSD: Gated Fusion Double SSD for Multispectral Pedestrian Detection
- FastPose: Towards Real-time Pose Estimation and Tracking via Scale-normalized Multi-task Networks
- Hallucinated-IQA: No-Reference Image Quality Assessment via Adversarial Learning
- Learning Better Features for Face Detection with Feature Fusion and Segmentation Supervision
- An Analysis of Pre-Training on Object Detection
- SFD: Single Shot Scale-invariant Face Detector
- VarifocalNet: An IoU-aware Dense Object Detector
- Improved Selective Refinement Network for Face Detection
- Diversify and Match: A Domain Adaptive Representation Learning Paradigm for Object Detection
- Object Detection in Video with Spatiotemporal Sampling Networks
- Looking Fast and Slow: Memory-Guided Mobile Video Object Detection
- Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd
- SSAP: Single-Shot Instance Segmentation With Affinity Pyramid
- Self-Training and Adversarial Background Regularization for Unsupervised Domain Adaptive One-Stage Object Detection
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Relation Distillation Networks for Video Object Detection
- Transferable Adversarial Attacks for Image and Video Object Detection
- An Annotation Saved is an Annotation Earned: Using Fully Synthetic Training for Object Instance Detection
- Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
- PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection
- Impression Network for Video Object Detection
- Exploring Categorical Regularization for Domain Adaptive Object Detection
- Scale-Aware Face Detection
- MegDet: A Large Mini-Batch Object Detector
- SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation
- Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey
- Faster RER-CNN: application to the detection of vehicles in aerial images
- Differentiating Objects by Motion: Joint Detection and Tracking of Small Flying Objects
- Improving Object Detection from Scratch via Gated Feature Reuse
- Robust Classification with Convolutional Prototype Learning
- Learning to Detect Human-Object Interactions
- Real-Time and Accurate Object Detection in Compressed Video by Long Short-term Feature Aggregation
- Instant-Teaching: An End-to-End Semi-Supervised Object Detection Framework
- DBF: Dynamic Belief Fusion for Combining Multiple Object Detectors
- ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations
- Progressive Sparse Local Attention for Video object detection
- Recent Advances in Deep Learning for Object Detection
- Corner Proposal Network for Anchor-free, Two-stage Object Detection
- Spatially Adaptive Computation Time for Residual Networks
- Extended Feature Pyramid Network for Small Object Detection
- Incremental Deep Learning for Robust Object Detection in Unknown Cluttered Environments
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- Multi-Object Tracking with Siamese Track-RCNN
- Dynamic Anchor Learning for Arbitrary-Oriented Object Detection
- Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships
- Towards Balanced Learning for Instance Recognition
- Adaptive NMS: Refining Pedestrian Detection in a Crowd
- Deep Contextual Attention for Human-Object Interaction Detection
- Accurate Face Detection for High Performance
- Robust and High Performance Face Detector
- Semantic Relation Reasoning for Shot-Stable Few-Shot Object Detection
- Repulsion Loss: Detecting Pedestrians in a Crowd
- Towards Unified INT8 Training for Convolutional Neural Network
- Sequential Context Encoding for Duplicate Removal
- Side-Aware Boundary Localization for More Precise Object Detection
- PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN
- Intrinsic Relationship Reasoning for Small Object Detection
- Towards Universal Object Detection by Domain Attention
- Implicit Feature Pyramid Network for Object Detection
- WSOD^2: Learning Bottom-up and Top-down Objectness Distillation for Weakly-supervised Object Detection
- Consistent Optimization for Single-Shot Object Detection
- Joint Detection and Tracking in Videos with Identification Features
- Rotation-Sensitive Regression for Oriented Scene Text Detection
- A Delay Metric for Video Object Detection: What Average Precision Fails to Tell
- Dense Label Encoding for Boundary Discontinuity Free Rotation Detection
- Improved Person Detection on Omnidirectional Images with Non-maxima Suppression
- Transformer-based Context Condensation for Boosting Feature Pyramids in Object Detection
- Learning Discriminative Motion Features Through Detection
- DeFRCN: Decoupled Faster R-CNN for Few-Shot Object Detection
- Feature Enhancement Network: A Refined Scene Text Detector
- PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
- Metamorphic Testing for Object Detection Systems
- Boundary Proposal Network for Two-Stage Natural Language Video Localization
- The Good, the Bad and the Ugly: Evaluating Convolutional Neural Networks for Prohibited Item Detection Using Real and Synthetically Composited X-ray Imagery
- Seeing Small Faces from Robust Anchor's Perspective
- Few-shot Object Detection via Feature Reweighting
- ACDnet: An action detection network for real-time edge computing based on flow-guided feature approximation and memory aggregation
- RoIMix: Proposal-Fusion among Multiple Images for Underwater Object Detection
- Learning Temporal Pose Estimation from Sparsely-Labeled Videos
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Dense RepPoints: Representing Visual Objects with Dense Point Sets
- RDSNet: A New Deep Architecture for Reciprocal Object Detection and Instance Segmentation
- Seeing isn't Believing: Practical Adversarial Attack Against Object Detectors
- Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition
- MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation
- Hit-Detector: Hierarchical Trinity Architecture Search for Object Detection
- RetinaTrack: Online Single Stage Joint Detection and Tracking
- BorderDet: Border Feature for Dense Object Detection
- A deep learning based solution for construction equipment detection: from development to deployment
- MoBiNet: A Mobile Binary Network for Image Classification
- C-RPNs: Promoting Object Detection in real world via a Cascade Structure of Region Proposal Networks
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- Memory Warps for Learning Long-Term Online Video Representations
- 3D Bounding Box Estimation for Autonomous Vehicles by Cascaded Geometric Constraints and Depurated 2D Detections Using 3D Results
- AdaZoom: Adaptive Zoom Network for Multi-Scale Object Detection in Large Scenes
- M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network
- Feature Intertwiner for Object Detection
- Clustered Object Detection in Aerial Images
- Accurate RGB-D Salient Object Detection via Collaborative Learning
- Learning Efficient Detector with Semi-supervised Adaptive Distillation
- IoU-aware Single-stage Object Detector for Accurate Localization
- Single-Shot Object Detection with Enriched Semantics
- Visual Relationship Detection using Scene Graphs: A Survey
- Higher-order Weighted Graph Convolutional Networks
- An Analysis of Scale Invariance in Object Detection - SNIP
- ICDAR 2019 Competition on Large-scale Street View Text with Partial Labeling -- RRC-LSVT
- Boundary-preserving Mask R-CNN
- MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object Detection
- Online Multiple Pedestrians Tracking using Deep Temporal Appearance Matching Association
- Multi-scale Location-aware Kernel Representation for Object Detection
- Learning a Layout Transfer Network for Context Aware Object Detection
- Soft Anchor-Point Object Detection
- Towards Interpretable R-CNN by Unfolding Latent Structures
- Single-Shot Bidirectional Pyramid Networks for High-Quality Object Detection
- Explicit Shape Encoding for Real-Time Instance Segmentation
- An Analysis of Deep Object Detectors For Diver Detection
- WordFence: Text Detection in Natural Images with Border Awareness
- Optimizing Video Object Detection via a Scale-Time Lattice
- Guided Attention Network for Object Detection and Counting on Drones
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- Learning Where to Focus for Efficient Video Object Detection
- Hetero-Center Loss for Cross-Modality Person Re-Identification
- End-to-End Wireframe Parsing
- Single Pixel Reconstruction for One-stage Instance Segmentation
- Mask R-CNN with Pyramid Attention Network for Scene Text Detection
- PanoNet: Real-time Panoptic Segmentation through Position-Sensitive Feature Embedding
- OMNIA Faster R-CNN: Detection in the wild through dataset merging and soft distillation
- A Hybrid Approach and Unified Framework for Bibliographic Reference Extraction
- RepGN:Object Detection with Relational Proposal Graph Network
- MFPN: A Novel Mixture Feature Pyramid Network of Multiple Architectures for Object Detection
- SpatialFlow: Bridging All Tasks for Panoptic Segmentation
- Toward Automatic Threat Recognition for Airport X-ray Baggage Screening with Deep Convolutional Object Detection
- Learning Region Features for Object Detection
- Multi-level Domain Adaptive learning for Cross-Domain Detection
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Crowd Counting Using Scale-Aware Attention Networks
- PolarMask++: Enhanced Polar Representation for Single-Shot Instance Segmentation and Beyond
- IoU-uniform R-CNN: Breaking Through the Limitations of RPN
- MnasFPN: Learning Latency-aware Pyramid Architecture for Object Detection on Mobile Devices
- Adversarial Feature Augmentation and Normalization for Visual Recognition
- DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation
- SMOT: Single-Shot Multi Object Tracking
- Rethinking Pseudo-LiDAR Representation
- Domain Adaptation from Synthesis to Reality in Single-model Detector for Video Smoke Detection
- TKD: Temporal Knowledge Distillation for Active Perception
- Dynamic Region-Aware Convolution
- Object Detection in Video with Spatial-temporal Context Aggregation
- Adaptive Feeding: Achieving Fast and Accurate Detections by Adaptively Combining Object Detectors
- Deep Feature Flow for Video Recognition
- MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
- Pose-based Modular Network for Human-Object Interaction Detection
- MSNet: A Multilevel Instance Segmentation Network for Natural Disaster Damage Assessment in Aerial Videos
- Accelerating Deep Learning Applications in Space
- End-To-End Face Detection and Recognition
- Weaving Multi-scale Context for Single Shot Detector
- Optimizing Region Selection for Weakly Supervised Object Detection
- Learning Spatial Pyramid Attentive Pooling in Image Synthesis and Image-to-Image Translation
- MultiResolution Attention Extractor for Small Object Detection
- SocialGuard: An Adversarial Example Based Privacy-Preserving Technique for Social Images
- Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning
- Fruit classification using deep feature maps in the presence of deceptive similar classes
- LapNet : Automatic Balanced Loss and Optimal Assignment for Real-Time Dense Object Detection
- ScratchDet: Training Single-Shot Object Detectors from Scratch
- DSFD: Dual Shot Face Detector
- Semi-convolutional Operators for Instance Segmentation
- An Overview Of 3D Object Detection
- Fast Object Detection in Compressed Video
- Deep Regionlets for Object Detection
- IPG-Net: Image Pyramid Guidance Network for Small Object Detection
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- Confluence: A Robust Non-IoU Alternative to Non-Maxima Suppression in Object Detection
- Robust Face Detection via Learning Small Faces on Hard Images
- ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
- RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection
- You Only Look One-level Feature
- Unmanned Aerial Vehicle Visual Detection and Tracking using Deep Neural Networks: A Performance Benchmark
- A Deep Ranking Model for Spatio-Temporal Highlight Detection from a 360 Video
- Optical Music Recognition: State of the Art and Major Challenges
- 3D Context Enhanced Region-based Convolutional Neural Network for End-to-End Lesion Detection
- Image Classification for Arabic: Assessing the Accuracy of Direct English to Arabic Translations
- Adaptive Object Detection with Dual Multi-Label Prediction
- Measuring economic activity from space: a case study using flying airplanes and COVID-19
- Towards High Performance Video Object Detection
- Recurrent Scale Approximation for Object Detection in CNN
- Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
- Weakly-Supervised Object Detection Learning through Human-Robot Interaction
- Detecting The Objects on The Road Using Modular Lightweight Network
- Road User Detection in Videos
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Learning Strict Identity Mappings in Deep Residual Networks
- 360-Indoor: Towards Learning Real-World Objects in 360° Indoor Equirectangular Images
- Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection
- Pseudo Mask Augmented Object Detection
- Towards Resolving the Challenge of Long-tail Distribution in UAV Images for Object Detection
- Pixel and Feature Level Based Domain Adaption for Object Detection in Autonomous Driving
- Image Cropping with Composition and Saliency Aware Aesthetic Score Map
- SiamMOT: Siamese Multi-Object Tracking
- Delving into Robust Object Detection from Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement Approach
- Characterizing Deep Learning Training Workloads on Alibaba-PAI
- DeRPN: Taking a further step toward more general object detection
- One-Shot Unsupervised Cross-Domain Detection
- V2F-Net: Explicit Decomposition of Occluded Pedestrian Detection
- Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection
- Rank of Experts: Detection Network Ensemble
- An Asymptotically Optimal Multi-Armed Bandit Algorithm and Hyperparameter Optimization
- SAN: Learning Relationship between Convolutional Features for Multi-Scale Object Detection
- Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person Search
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Towards Multi-class Object Detection in Unconstrained Remote Sensing Imagery
- Deep Learning Enhanced Extended Depth-of-Field for Thick Blood-Film Malaria High-Throughput Microscopy
- DuBox: No-Prior Box Objection Detection via Residual Dual Scale Detectors
- CCL: Cross-modal Correlation Learning with Multi-grained Fusion by Hierarchical Network
- Dynamic Edge Weights in Graph Neural Networks for 3D Object Detection
- AlphaRotate: A Rotation Detection Benchmark using TensorFlow
- Geometry-Aware Video Object Detection for Static Cameras
- Fast Region Proposal Learning for Object Detection for Robotics
- Ensemble Soft-Margin Softmax Loss for Image Classification
- Exploiting Web Images for Weakly Supervised Object Detection
- Precipitation Forecasting via Multi-Scale Deconstructed ConvLSTM
- Relational Learning for Joint Head and Human Detection
- Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object Detection
- Triply Supervised Decoder Networks for Joint Detection and Segmentation
- A Structured Model For Action Detection
- Adaptive Class Suppression Loss for Long-Tail Object Detection
- Object Detection based on Region Decomposition and Assembly
- Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes
- Re-ID Driven Localization Refinement for Person Search
- Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection
- Deep Regression Forests for Age Estimation
- Segmentation Mask Guided End-to-End Person Search
- An End-to-End Neural Network for Image Cropping by Learning Composition from Aesthetic Photos
- Real-Time Rotation-Invariant Face Detection with Progressive Calibration Networks
- Beyond Trade-off: Accelerate FCN-based Face Detector with Higher Accuracy
- OnlineAugment: Online Data Augmentation with Less Domain Knowledge
- NOH-NMS: Improving Pedestrian Detection by Nearby Objects Hallucination
- Straight to Shapes++: Real-time Instance Segmentation Made More Accurate
- Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
- Yes-Net: An effective Detector Based on Global Information
- GTNet: Generative Transfer Network for Zero-Shot Object Detection
- Propose-and-Attend Single Shot Detector
- Dynamic Filtering with Large Sampling Field for ConvNets
- TYolov5: A Temporal Yolov5 Detector Based on Quasi-Recurrent Neural Networks for Real-Time Handgun Detection in Video
- Scaling Object Detection by Transferring Classification Weights
- Object-aware Feature Aggregation for Video Object Detection
- User Constrained Thumbnail Generation using Adaptive Convolutions
- Rotated Feature Network for multi-orientation object detection
- GAN-Knowledge Distillation for one-stage Object Detection
- Pixelwise Instance Segmentation with a Dynamically Instantiated Network
- Reducing Label Noise in Anchor-Free Object Detection
- Towards Better Object Detection in Scale Variation with Adaptive Feature Selection
- TensorFlow with user friendly Graphical Framework for object detection API
- Large-scale image analysis using docker sandboxing
- Training Generative Adversarial Networks in One Stage
- How low can you go? Privacy-preserving people detection with an omni-directional camera
- Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild
- iShape: A First Step Towards Irregular Shape Instance Segmentation
- Exploring the Capacity of an Orderless Box Discretization Network for Multi-orientation Scene Text Detection
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Deep Neural Network Based Real-time Kiwi Fruit Flower Detection in an Orchard Environment
- In Defense of the Classification Loss for Person Re-Identification
- AFD-Net: Adaptive Fully-Dual Network for Few-Shot Object Detection
- Tell Me What They're Holding: Weakly-supervised Object Detection with Transferable Knowledge from Human-object Interaction
- Binge Watching: Scaling Affordance Learning from Sitcoms
- CPM R-CNN: Calibrating Point-guided Misalignment in Object Detection
- Emotion Generation and Recognition: A StarGAN Approach
- Detecting 11K Classes: Large Scale Object Detection without Fine-Grained Bounding Boxes
- Dually Supervised Feature Pyramid for Object Detection and Segmentation
- LiDAR and Camera Detection Fusion in a Real Time Industrial Multi-Sensor Collision Avoidance System
- Deep Learning Methods for Real-time Detection and Analysis of Wagner Ulcer Classification System
- Representation Sharing for Fast Object Detector Search and Beyond
- COBE: Contextualized Object Embeddings from Narrated Instructional Video
- ElixirNet: Relation-aware Network Architecture Adaptation for Medical Lesion Detection
- Online Anomaly Detection in Surveillance Videos with Asymptotic Bounds on False Alarm Rate
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- Inception Convolution with Efficient Dilation Search
- ACNet: Mask-Aware Attention with Dynamic Context Enhancement for Robust Acne Detection
- Finding a Needle in a Haystack: Tiny Flying Object Detection in 4K Videos using a Joint Detection-and-Tracking Approach
- Learning to Track Object Position through Occlusion
- High-speed Railway Fastener Detection and Localization Method based on convolutional neural network
- Image Captioning with Unseen Objects
- Object as Distribution
- Multilevel Knowledge Transfer for Cross-Domain Object Detection
- Patchwork: A Patch-wise Attention Network for Efficient Object Detection and Segmentation in Video Streams
- A Single-shot Object Detector with Feature Aggragation and Enhancement
- PatchNet -- Short-range Template Matching for Efficient Video Processing
- Detect or Track: Towards Cost-Effective Video Object Detection/Tracking
- Skeleton-Based Action Recognition with Synchronous Local and Non-local Spatio-temporal Learning and Frequency Attention
- Probabilistic Oriented Object Detection in Automotive Radar
- GraftNet: An Engineering Implementation of CNN for Fine-grained Multi-label Task
- Cognitive Deep Machine Can Train Itself
- StackNet: Stacking Parameters for Continual learning
- A study of the effect of the illumination model on the generation of synthetic training datasets
- Object Detection from Scratch with Deep Supervision
- AABO: Adaptive Anchor Box Optimization for Object Detection via Bayesian Sub-sampling
- HybridNet: Classification and Reconstruction Cooperation for Semi-Supervised Learning
- What leads to generalization of object proposals?
- Semi-Anchored Detector for One-Stage Object Detection
- Performance of object recognition in wearable videos
- SwiftFace: Real-Time Face Detection
- Scale Match for Tiny Person Detection
- S4Net: Single Stage Salient-Instance Segmentation
- Multiple receptive fields and small-object-focusing weakly-supervised segmentation network for fast object detection
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
- DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection under Partial Occlusion
- Evaluation of Model Selection for Kernel Fragment Recognition in Corn Silage
- PointINS: Point-based Instance Segmentation
- Localize to Classify and Classify to Localize: Mutual Guidance in Object Detection
- Multi-frame Collaboration for Effective Endoscopic Video Polyp Detection via Spatial-Temporal Feature Transformation
- Learning Instance-Aware Object Detection Using Determinantal Point Processes
- EHSOD: CAM-Guided End-to-end Hybrid-Supervised Object Detection with Cascade Refinement
- FA-RPN: Floating Region Proposals for Face Detection
- ASSD: Attentive Single Shot Multibox Detector
- Quantization Mimic: Towards Very Tiny CNN for Object Detection
- Precise Box Score: Extract More Information from Datasets to Improve the Performance of Face Detection
- Improving Object Detection with Selective Self-supervised Self-training
- Extract and Merge: Merging extracted humans from different images utilizing Mask R-CNN
- LiDAR-assisted Large-scale Privacy Protection in Street-view Cycloramas
- LPM: Learnable Pooling Module for Efficient Full-Face Gaze Estimation
- Optimizing Data Processing in Space for Object Detection in Satellite Imagery
- Digging Deeper into CRNN Model in Chinese Text Images Recognition
- Stingray Detection of Aerial Images Using Augmented Training Images Generated by A Conditional Generative Model
- Natural Language Video Localization with Learnable Moment Proposals
- Joint Distribution Alignment via Adversarial Learning for Domain Adaptive Object Detection
- Machine Vision in the Context of Robotics: A Systematic Literature Review
- Creating Lightweight Object Detectors with Model Compression for Deployment on Edge Devices
- Semantic Image Retrieval by Uniting Deep Neural Networks and Cognitive Architectures
- Image Matters: Scalable Detection of Offensive and Non-Compliant Content / Logo in Product Images
- Deep Learning Acceleration Techniques for Real Time Mobile Vision Applications
- RUHSNet: 3D Object Detection Using Lidar Data in Real Time
- FPAN: Fine-grained and Progressive Attention Localization Network for Data Retrieval
- Open-World Entity Segmentation
- Enabling Incremental Knowledge Transfer for Object Detection at the Edge
- TLGAN: document Text Localization using Generative Adversarial Nets
- G-RCN: Optimizing the Gap between Classification and Localization Tasks for Object Detection
- Bio-Inspired Adversarial Attack Against Deep Neural Networks
- Learning Polar Encodings for Arbitrary-Oriented Ship Detection in SAR Images
- Semantic Segmentation via Highly Fused Convolutional Network with Multiple Soft Cost Functions
- An Action Recognition network for specific target based on rMC and RPN
- Feature Selective Networks for Object Detection
- Multi Target Tracking by Learning from Generalized Graph Differences
- Robust Handwriting Recognition with Limited and Noisy Data
- Visible Feature Guidance for Crowd Pedestrian Detection
- ShuffleDet: Real-Time Vehicle Detection Network in On-board Embedded UAV Imagery
- Minimizing Supervision in Multi-label Categorization
- A Novel Integrated Framework for Learning both Text Detection and Recognition
- Which to Match? Selecting Consistent GT-Proposal Assignment for Pedestrian Detection
- Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation
- PBRnet: Pyramidal Bounding Box Refinement to Improve Object Localization Accuracy
- Fusion of an Ensemble of Augmented Image Detectors for Robust Object Detection
- Revisit Multinomial Logistic Regression in Deep Learning: Data Dependent Model Initialization for Image Recognition
- Improve CAM with Auto-adapted Segmentation and Co-supervised Augmentation
- What Can Help Pedestrian Detection?
- Feature Flow: In-network Feature Flow Estimation for Video Object Detection
- Unsupervised Spiking Instance Segmentation on Event Data using STDP
- Object Detection in Specific Traffic Scenes using YOLOv2
- MetricUNet: Synergistic Image- and Voxel-Level Learning for Precise CT Prostate Segmentation via Online Sampling
- Impoved RPN for Single Targets Detection based on the Anchor Mask Net
- OICSR: Out-In-Channel Sparsity Regularization for Compact Deep Neural Networks
- Higher-order Network for Action Recognition
- BAN: Focusing on Boundary Context for Object Detection
- Bridging the Gap Between Object Detection and User Intent via Query-Modulation
- Condensing Two-stage Detection with Automatic Object Key Part Discovery
- Human Object Interaction Detection using Two-Direction Spatial Enhancement and Exclusive Object Prior
- Ventral-Dorsal Neural Networks: Object Detection via Selective Attention
- KL-Divergence-Based Region Proposal Network for Object Detection
- Vision-based Price Suggestion for Online Second-hand Items
- End-to-end Learning of Image based Lane-Change Decision
- Unsupervised Part Discovery via Feature Alignment
- Learning Sports Camera Selection from Internet Videos
- Localizing Grouped Instances for Efficient Detection in Low-Resource Scenarios
- FPAENet: Pneumonia Detection Network Based on Feature Pyramid Attention Enhancement
- S-OHEM: Stratified Online Hard Example Mining for Object Detection
- DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object Detection
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- CNN-based Human Detection for UAVs in Search and Rescue
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- Few-Shot Video Object Detection
- ARMA Nets: Expanding Receptive Field for Dense Prediction
- Temporal RoI Align for Video Object Recognition
- Densely Semantic Enhancement for Domain Adaptive Region-free Detectors
- EOLO: Embedded Object Segmentation only Look Once
- FSD: Feature Skyscraper Detector for Stem End and Blossom End of Navel Orange
- Vehicle Image Generation Going Well with The Surroundings
- CS-R-FCN: Cross-supervised Learning for Large-Scale Object Detection
- Two-stage multi-scale breast mass segmentation for full mammogram analysis without user intervention
- Layer-wise Customized Weak Segmentation Block and AIoU Loss for Accurate Object Detection
- Log-Polar Space Convolution for Convolutional Neural Networks
- Modification method for single-stage object detectors that allows to exploit the temporal behaviour of a scene to improve detection accuracy
- RoIFusion: 3D Object Detection from LiDAR and Vision
- Soccer on Your Tabletop
- Perceiving Traffic from Aerial Images
- External-Memory Networks for Low-Shot Learning of Targets in Forward-Looking-Sonar Imagery
- Semantically-Aware Strategies for Stereo-Visual Robotic Obstacle Avoidance
- Feature Fusion Detector for Semantic Cognition of Remote Sensing
- FPGA Implementations of 3D-SIMD Processor Architecture for Deep Neural Networks Using Relative Indexed Compressed Sparse Filter Encoding Format and Stacked Filters Stationary Flow
- ScarfNet: Multi-scale Features with Deeply Fused and Redistributed Semantics for Enhanced Object Detection
- Integrating Deep Learning and Augmented Reality to Enhance Situational Awareness in Firefighting Environments
- DupNet: Towards Very Tiny Quantized CNN with Improved Accuracy for Face Detection
- PosNeg-Balanced Anchors with Aligned Features for Single-Shot Object Detection
- Inspect Transfer Learning Architecture with Dilated Convolution
- PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS with Relationship Recovery
- Model Adaption Object Detection System for Robot
- AGSFCOS: Based on attention mechanism and Scale-Equalizing pyramid network of object detection
- Occluded Pedestrian Detection with Visible IoU and Box Sign Predictor
- Minimum Delay Object Detection From Video
- Human-Object Interaction Detection via Weak Supervision
- Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning
- Universal Physical Camouflage Attacks on Object Detectors
- A Machine Learning Framework for Data Ingestion in Document Images
- Advancement of Deep Learning in Pneumonia and Covid-19 Classification and Localization: A Qualitative and Quantitative Analysis
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- A Top-down Approach to Articulated Human Pose Estimation and Tracking
- Large-scale mammography CAD with Deformable Conv-Nets
- Detector-in-Detector: Multi-Level Analysis for Human-Parts
- Detecting Lesion Bounding Ellipses With Gaussian Proposal Networks
- U-net super-neural segmentation and similarity calculation to realize vegetation change assessment in satellite imagery