Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
arXiv:1406.4729 · doi:10.1007/978-3-319-10578-9_23
Abstract
Existing deep convolutional neural networks (CNNs) require a fixed-size (e.g., 224x224) input image. This requirement is "artificial" and may reduce the recognition accuracy for the images or sub-images of an arbitrary size/scale. In this work, we equip the networks with another pooling strategy, "spatial pyramid pooling", to eliminate the above requirement. The new network structure, called SPP-net, can generate a fixed-length representation regardless of image size/scale. Pyramid pooling is also robust to object deformations. With these advantages, SPP-net should in general improve all CNN-based image classification methods. On the ImageNet 2012 dataset, we demonstrate that SPP-net boosts the accuracy of a variety of CNN architectures despite their different designs. On the Pascal VOC 2007 and Caltech101 datasets, SPP-net achieves state-of-the-art classification results using a single full-image representation and no fine-tuning. The power of SPP-net is also significant in object detection. Using SPP-net, we compute the feature maps from the entire image only once, and then pool features in arbitrary regions (sub-images) to generate fixed-length representations for training the detectors. This method avoids repeatedly computing the convolutional features. In processing test images, our method is 24-102x faster than the R-CNN method, while achieving better or comparable accuracy on Pascal VOC 2007. In ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2014, our methods rank #2 in object detection and #3 in image classification among all 38 teams. This manuscript also introduces the improvement made for this competition.
This manuscript is the accepted version for IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2015. See Changelog
References in corpus (4)
Cited by in corpus (408)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- SSD: Single Shot MultiBox Detector
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Deep Residual Learning for Image Recognition
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
- Attention Mechanisms in Computer Vision: A Survey
- Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
- Empirical Evaluation of Rectified Activations in Convolutional Network
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- Focal Loss for Dense Object Detection
- Spatio-Temporal Backpropagation for Training High-performance Spiking Neural Networks
- ParseNet: Looking Wider to See Better
- CNN: Single-label to Multi-label
- Learning to Segment Object Candidates
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Fixed Point Quantization of Deep Convolutional Networks
- Online Tracking by Learning Discriminative Saliency Map with Convolutional Neural Network
- It's Written All Over Your Face: Full-Face Appearance-Based Gaze Estimation
- Convolutional Feature Masking for Joint Object and Stuff Segmentation
- Feature Pyramid Networks for Object Detection
- DenseBox: Unifying Landmark Localization with End to End Object Detection
- Deformable Convolutional Networks
- Texture Synthesis Using Convolutional Neural Networks
- Visual Saliency Detection Based on Multiscale Deep CNN Features
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Beyond Skip Connections: Top-Down Modulation for Object Detection
- Scalable, High-Quality Object Detection
- Systematic evaluation of CNN advances on the ImageNet
- Compute Trends Across Three Eras of Machine Learning
- A Review on Deep Learning Techniques for Video Prediction
- Pyramid Scene Parsing Network
- Semantic Labeling in Very High Resolution Images via a Self-Cascaded Convolutional Neural Network
- Transferring Rich Feature Hierarchies for Robust Visual Tracking
- Loss Functions and Metrics in Deep Learning
- A Review of Object Detection Models based on Convolutional Neural Network
- SlimYOLOv3: Narrower, Faster and Better for Real-Time UAV Applications
- Universal Correspondence Network
- R-CNN: Fast Tiny Object Detection in Large-Scale Remote Sensing Images
- Image De-raining Using a Conditional Generative Adversarial Network
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Deep Learning for Generic Object Detection: A Survey
- Detection and Localization of Robotic Tools in Robot-Assisted Surgery Videos Using Deep Neural Networks for Region Proposal and Detection
- Improving Object Detection with Deep Convolutional Networks via Bayesian Optimization and Structured Prediction
- Learning to Compare Image Patches via Convolutional Neural Networks
- LSDA: Large Scale Detection Through Adaptation
- Image Forgery Localization Based on Multi-Scale Convolutional Neural Networks
- On Low-Resolution Face Recognition in the Wild: Comparisons and New Techniques
- Position Detection and Direction Prediction for Arbitrary-Oriented Ships via Multitask Rotation Region Convolutional Neural Network
- Learning a Rotation Invariant Detector with Rotatable Bounding Box
- Recent Advance in Content-based Image Retrieval: A Literature Survey
- End-to-End Photo-Sketch Generation via Fully Convolutional Representation Learning
- Knowledge Guided Disambiguation for Large-Scale Scene Classification with Multi-Resolution CNNs
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
- Hydra: an Ensemble of Convolutional Neural Networks for Geospatial Land Classification
- Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- CornerNet: Detecting Objects as Paired Keypoints
- SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
- DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic
- Image Super-Resolution Using Deep Convolutional Networks
- Temporal Action Detection with Structured Segment Networks
- Joint Unsupervised Learning of Deep Representations and Image Clusters
- A deep learning framework for quality assessment and restoration in video endoscopy
- X-Net: Brain Stroke Lesion Segmentation Based on Depthwise Separable Convolution and Long-range Dependencies
- Libra R-CNN: Towards Balanced Learning for Object Detection
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- DeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection
- Deep Contrast Learning for Salient Object Detection
- Good Practice in CNN Feature Transfer
- Instance-Aware Hashing for Multi-Label Image Retrieval
- Learning Contextual Dependencies with Convolutional Hierarchical Recurrent Neural Networks
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- A Discriminative Representation of Convolutional Features for Indoor Scene Recognition
- Self Paced Deep Learning for Weakly Supervised Object Detection
- Semantics-Guided Contrastive Network for Zero-Shot Object detection
- Scale-aware Fast R-CNN for Pedestrian Detection
- Generalizing Pooling Functions in Convolutional Neural Networks: Mixed, Gated, and Tree
- Single-Shot Refinement Neural Network for Object Detection
- Look Wider to Match Image Patches with Convolutional Neural Networks
- Bottom-up Object Detection by Grouping Extreme and Center Points
- ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
- Flow-Guided Feature Aggregation for Video Object Detection
- RCS-YOLO: A Fast and High-Accuracy Object Detector for Brain Tumor Detection
- DeepText: A Unified Framework for Text Proposal Generation and Text Detection in Natural Images
- Deep Learning Face Attributes in the Wild
- Cascade R-CNN: High Quality Object Detection and Instance Segmentation
- What Do We Understand About Convolutional Networks?
- Rethinking ImageNet Pre-training
- Density-aware Single Image De-raining using a Multi-stream Dense Network
- Attention-based Pyramid Aggregation Network for Visual Place Recognition
- LittleYOLO-SPP: A Delicate Real-Time Vehicle Detection Algorithm
- Learning a Discriminative Feature Network for Semantic Segmentation
- Too Far to See? Not Really! --- Pedestrian Detection with Scale-aware Localization Policy
- Deep Image Retrieval: Learning global representations for image search
- Instance-aware Semantic Segmentation via Multi-task Network Cascades
- ME R-CNN: Multi-Expert R-CNN for Object Detection
- Exploiting Image-trained CNN Architectures for Unconstrained Video Classification
- DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection
- Object Detection Through Exploration With A Foveated Visual Field
- Object Detection Using Sim2Real Domain Randomization for Robotic Applications
- Multi-Class Multi-Object Tracking using Changing Point Detection
- Contextual Action Recognition with R*CNN
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Boundary-Aware Segmentation Network for Mobile and Web Applications
- Pedestrian-Synthesis-GAN: Generating Pedestrian Data in Real Scene and Beyond
- Relation Networks for Object Detection
- Fixed-sized representation learning from Offline Handwritten Signatures of different sizes
- RON: Reverse Connection with Objectness Prior Networks for Object Detection
- Pyramid Stereo Matching Network
- Stacking-Based Deep Neural Network: Deep Analytic Network for Pattern Classification
- LOGO-Net: Large-scale Deep Logo Detection and Brand Recognition with Deep Region-based Convolutional Networks
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Learning to Refine Object Segments
- Poisson CNN: Convolutional neural networks for the solution of the Poisson equation on a Cartesian mesh
- Detecting and Recognizing Human-Object Interactions
- Horizontal Pyramid Matching for Person Re-identification
- Analysis and Optimization of Convolutional Neural Network Architectures
- Multimodal Convolutional Neural Networks for Matching Image and Sentence
- Learning Complexity-Aware Cascades for Deep Pedestrian Detection
- SNIPER: Efficient Multi-Scale Training
- All Grains, One Scheme (AGOS): Learning Multi-grain Instance Representation for Aerial Scene Classification
- Waterfall Atrous Spatial Pooling Architecture for Efficient Semantic Segmentation
- End-to-End Image Super-Resolution via Deep and Shallow Convolutional Networks
- Visualizing and Comparing Convolutional Neural Networks
- Search and Rescue with Airborne Optical Sectioning
- Knowledge-aware Deep Framework for Collaborative Skin Lesion Segmentation and Melanoma Recognition
- Range Loss for Deep Face Recognition with Long-tail
- Video Semantic Segmentation with Distortion-Aware Feature Correction
- Learning Effective RGB-D Representations for Scene Recognition
- Deep convolutional filter banks for texture recognition and segmentation
- AMNet: Deep Atrous Multiscale Stereo Disparity Estimation Networks
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
- Unsupervised Semantic-based Aggregation of Deep Convolutional Features
- ResFeats: Residual Network Based Features for Image Classification
- Discriminative out-of-distribution detection for semantic segmentation
- Unsupervised Adaptation for Synthetic-to-Real Handwritten Word Recognition
- Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
- ThunderNet: Towards Real-time Generic Object Detection
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- A deep architecture for unified aesthetic prediction
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Deep Spatial Pyramid: The Devil is Once Again in the Details
- Deep Stacked Hierarchical Multi-patch Network for Image Deblurring
- Learning based Facial Image Compression with Semantic Fidelity Metric
- HD-CNN: Hierarchical Deep Convolutional Neural Network for Large Scale Visual Recognition
- Rethinking Classification and Localization for Object Detection
- Photo Aesthetics Ranking Network with Attributes and Content Adaptation
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- SPP-Net: Deep Absolute Pose Regression with Synthetic Views
- Deep Learning For Computer Vision Tasks: A review
- Identifying Generalization Properties in Neural Networks
- Monocular 3D Object Detection with Sequential Feature Association and Depth Hint Augmentation
- Weakly- and Semi-Supervised Object Detection with Expectation-Maximization Algorithm
- Multi-Task Learning for Left Atrial Segmentation on GE-MRI
- Spatial-temporal Conv-sequence Learning with Accident Encoding for Traffic Flow Prediction
- Towards a Visual Turing Challenge
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
- Hybrid Channel Based Pedestrian Detection
- CyCNN: A Rotation Invariant CNN using Polar Mapping and Cylindrical Convolution Layers
- Fisher Kernel for Deep Neural Activations
- CST-YOLO: A Novel Method for Blood Cell Detection Based on Improved YOLOv7 and CNN-Swin Transformer
- DeepBox: Learning Objectness with Convolutional Networks
- Learning Deep ResNet Blocks Sequentially using Boosting Theory
- Accelerating Very Deep Convolutional Networks for Classification and Detection
- Semantic segmentation of mFISH images using convolutional networks
- Looking Fast and Slow: Memory-Guided Mobile Video Object Detection
- Co-salient Object Detection Based on Deep Saliency Networks and Seed Propagation over an Integrated Graph
- Scene Graph Generation from Objects, Phrases and Region Captions
- Impression Network for Video Object Detection
- A Discriminative CNN Video Representation for Event Detection
- MegDet: A Large Mini-Batch Object Detector
- Convolutional Neural Networks at Constrained Time Cost
- Target Detection and Segmentation in Circular-Scan Synthetic-Aperture-Sonar Images using Semi-Supervised Convolutional Encoder-Decoders
- Weakly Supervised Dense Video Captioning
- Initialization Strategies of Spatio-Temporal Convolutional Neural Networks
- Effect of Annotation Errors on Drone Detection with YOLOv3
- Deep Neural Networks for Marine Debris Detection in Sonar Images
- Progressive Sparse Local Attention for Video object detection
- Zoom Out-and-In Network with Recursive Training for Object Proposal
- Resolution learning in deep convolutional networks using scale-space theory
- Detection in Crowded Scenes: One Proposal, Multiple Predictions
- Crafting GBD-Net for Object Detection
- Towards Balanced Learning for Instance Recognition
- Stacked Pooling: Improving Crowd Counting by Boosting Scale Invariance
- Scale-aware Pixel-wise Object Proposal Networks
- Residual Features and Unified Prediction Network for Single Stage Detection
- Cross Modal Distillation for Supervision Transfer
- Fast Low-rank Representation based Spatial Pyramid Matching for Image Classification
- An End-to-End Network for Panoptic Segmentation
- Detecting Small Objects in Thermal Images Using Single-Shot Detector
- Coarse to Fine Multi-Resolution Temporal Convolutional Network
- Network of Experts for Large-Scale Image Categorization
- Increasing the Robustness of Semantic Segmentation Models with Painting-by-Numbers
- Hypercolumns for Object Segmentation and Fine-grained Localization
- Boosting Convolutional Features for Robust Object Proposals
- Co-localization with Category-Consistent Features and Geodesic Distance Propagation
- SPP-CNN: An Efficient Framework for Network Robustness Prediction
- Sequential Context Encoding for Duplicate Removal
- A new take on measuring relative nutritional density: The feasibility of using a deep neural network to assess commercially-prepared pureed food concentrations
- Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data
- A Unified Deep Learning Framework for Short-Duration Speaker Verification in Adverse Environments
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- DeePM: A Deep Part-Based Model for Object Detection and Semantic Part Localization
- Consistent Optimization for Single-Shot Object Detection
- Selective Unsupervised Feature Learning with Convolutional Neural Network (S-CNN)
- Iterative Instance Segmentation
- On The Stability of Video Detection and Tracking
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
- Learning Discriminative Motion Features Through Detection
- Crafting a multi-task CNN for viewpoint estimation
- Spatially-Adaptive Filter Units for Compact and Efficient Deep Neural Networks
- Adaptive Object Detection Using Adjacency and Zoom Prediction
- Learning Temporal Pose Estimation from Sparsely-Labeled Videos
- Few-shot Object Detection via Feature Reweighting
- Efficient Object Detection for High Resolution Images
- Solution for Large-Scale Hierarchical Object Detection Datasets with Incomplete Annotation and Data Imbalance
- Crowd Counting via Weighted VLAD on Dense Attribute Feature Maps
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Convolutional Neural Networks for joint object detection and pose estimation: A comparative study
- Cascaded Structure Tensor Framework for Robust Identification of Heavily Occluded Baggage Items from X-ray Scans
- Modeling Visual Compatibility through Hierarchical Mid-level Elements
- BorderDet: Border Feature for Dense Object Detection
- C-RPNs: Promoting Object Detection in real world via a Cascade Structure of Region Proposal Networks
- Enhancing sea ice segmentation in Sentinel-1 images with atrous convolutions
- DSNet for Real-Time Driving Scene Semantic Segmentation
- An Analysis of Scale Invariance in Object Detection - SNIP
- Improved Part Segmentation Performance by Optimising Realism of Synthetic Images using Cycle Generative Adversarial Networks
- Active Convolution: Learning the Shape of Convolution for Image Classification
- Towards Asteroid Detection in Microlensing Surveys with Deep Learning
- A HMAX with LLC for visual recognition
- Learning Representative Temporal Features for Action Recognition
- Learning a Layout Transfer Network for Context Aware Object Detection
- Weakly-supervised Learning of Mid-level Features for Pedestrian Attribute Recognition and Localization
- AttentionNet: Aggregating Weak Directions for Accurate Object Detection
- Rethinking the backbone architecture for tiny object detection
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Automatic Detection of Interplanetary Coronal Mass Ejections in Solar Wind In Situ Data
- Adaptive Fractional Dilated Convolution Network for Image Aesthetics Assessment
- Object Discovery via Cohesion Measurement
- Person Search by Multi-Scale Matching
- A Jointly Learned Deep Architecture for Facial Attribute Analysis and Face Detection in the Wild
- Improving Deep Neural Network with Multiple Parametric Exponential Linear Units
- Collaborative Learning for Weakly Supervised Object Detection
- Tackling Catastrophic Forgetting and Background Shift in Continual Semantic Segmentation
- Semiotic Aggregation in Deep Learning
- Learning Region Features for Object Detection
- Generating Hard Examples for Pixel-wise Classification
- RepGN:Object Detection with Relational Proposal Graph Network
- Learning Where to Focus for Efficient Video Object Detection
- Hard-Attention for Scalable Image Classification
- Cascaded Structure Tensor Framework for Robust Identification of Heavily Occluded Baggage Items from Multi-Vendor X-ray Scans
- Face Detection with Feature Pyramids and Landmarks
- Deep TEN: Texture Encoding Network
- Deep Co-attention based Comparators For Relative Representation Learning in Person Re-identification
- What Is the Best Practice for CNNs Applied to Visual Instance Retrieval?
- Deep Feature Flow for Video Recognition
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- Integrated perception with recurrent multi-task neural networks
- TriLiteNet: Lightweight Model for Multi-Task Visual Perception
- Deep Rigid Instance Scene Flow
- Window-Object Relationship Guided Representation Learning for Generic Object Detections
- SampleAhead: Online Classifier-Sampler Communication for Learning from Synthesized Data
- Using Depth for Pixel-Wise Detection of Adversarial Attacks in Crowd Counting
- Is Faster R-CNN Doing Well for Pedestrian Detection?
- Deep Regionlets for Object Detection
- Learning Fine-grained Features via a CNN Tree for Large-scale Classification
- Recovering hard-to-find object instances by sampling context-based object proposals
- Do More Dropouts in Pool5 Feature Maps for Better Object Detection
- Displacement-Invariant Cost Computation for Efficient Stereo Matching
- Deep Roto-Translation Scattering for Object Classification
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- An Out-of-the-box Full-network Embedding for Convolutional Neural Networks
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Reflective Decoding Network for Image Captioning
- Face Detection, Bounding Box Aggregation and Pose Estimation for Robust Facial Landmark Localisation in the Wild
- Context-Aware Crowd Counting
- Camera View Adjustment Prediction for Improving Image Composition
- Tag Prediction at Flickr: a View from the Darkroom
- Factors in Finetuning Deep Model for object detection
- Mid-level Elements for Object Detection
- Towards High Performance Video Object Detection
- Where to Focus: Deep Attention-based Spatially Recurrent Bilinear Networks for Fine-Grained Visual Recognition
- Towards Precise End-to-end Weakly Supervised Object Detection Network
- Robust Real-time Pedestrian Detection in Aerial Imagery on Jetson TX2
- Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
- An Effective Two-Branch Model-Based Deep Network for Single Image Deraining
- Detect-and-describe: Joint learning framework for detection and description of objects
- Cascaded Subpatch Networks for Effective CNNs
- Learning Discriminative Features via Label Consistent Neural Network
- CathAI: Fully Automated Interpretation of Coronary Angiograms Using Neural Networks
- Streaming egocentric action anticipation: An evaluation scheme and approach
- Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
- Delving into Robust Object Detection from Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement Approach
- Active Fine-Tuning from gMAD Examples Improves Blind Image Quality Assessment
- PMC-GANs: Generating Multi-Scale High-Quality Pedestrian with Multimodal Cascaded GANs
- Image Recognition Using Scale Recurrent Neural Networks
- Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection
- Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object Detection
- Farm land weed detection with region-based deep convolutional neural networks
- Learning to Segment Moving Objects in Videos
- GLSD: The Global Large-Scale Ship Database and Baseline Evaluations
- Detecting Small, Densely Distributed Objects with Filter-Amplifier Networks and Loss Boosting
- PCC Net: Perspective Crowd Counting via Spatial Convolutional Network
- SeGAN: Segmenting and Generating the Invisible
- Triply Supervised Decoder Networks for Joint Detection and Segmentation
- Superpixel Convolutional Networks using Bilateral Inceptions
- DeepSFM: Structure From Motion Via Deep Bundle Adjustment
- Edge-Cloud Collaborated Object Detection via Difficult-Case Discriminator
- Inability of spatial transformations of CNN feature maps to support invariant recognition
- FarSee-Net: Real-Time Semantic Segmentation by Efficient Multi-scale Context Aggregation and Feature Space Super-resolution
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- Trapped in texture bias? A large scale comparison of deep instance segmentation
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- Autonomous Navigation in Dynamic Environments: Deep Learning-Based Approach
- Unit panel nodes detection by CNN on FAST reflector
- Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019
- Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo Classification
- Contrast-Oriented Deep Neural Networks for Salient Object Detection
- Translate-to-Recognize Networks for RGB-D Scene Recognition
- ViP-CNN: Visual Phrase Guided Convolutional Neural Network
- A Taught-Obesrve-Ask (TOA) Method for Object Detection with Critical Supervision
- Multiplier-less Artificial Neurons Exploiting Error Resiliency for Energy-Efficient Neural Computing
- Learning Video-Story Composition via Recurrent Neural Network
- Large-scale image analysis using docker sandboxing
- The Treasure beneath Convolutional Layers: Cross-convolutional-layer Pooling for Image Classification
- S3Pool: Pooling with Stochastic Spatial Sampling
- Inferring Distributions Over Depth from a Single Image
- COBE: Contextualized Object Embeddings from Narrated Instructional Video
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
- Visual Recognition Using Directional Distribution Distance
- A Survey of FPGA-Based Robotic Computing
- Using Deep Networks for Drone Detection
- Spatially-Adaptive Filter Units for Deep Neural Networks
- Parsing R-CNN for Instance-Level Human Analysis
- Diagnosing State-Of-The-Art Object Proposal Methods
- Convolutional Tables Ensemble: classification in microseconds
- A Deep Learning Approach to Drone Monitoring
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Optimising the Input Image to Improve Visual Relationship Detection
- CamLoc: Pedestrian Location Detection from Pose Estimation on Resource-constrained Smart-cameras
- Object 6D Pose Estimation with Non-local Attention
- ASSD: Attentive Single Shot Multibox Detector
- Convolution in Convolution for Network in Network
- Dual Temporal Memory Network for Efficient Video Object Segmentation
- Feature Selective Networks for Object Detection
- Quantization Mimic: Towards Very Tiny CNN for Object Detection
- Driving among Flatmobiles: Bird-Eye-View occupancy grids from a monocular camera for holistic trajectory planning
- Convolutional module for heart localization and segmentation in MRI
- Detecting Temporally Consistent Objects in Videos through Object Class Label Propagation
- Mask R-CNN Based Object Detection for Intelligent Wireless Power Transfer
- Towards Locally Consistent Object Counting with Constrained Multi-stage Convolutional Neural Networks
- The Devils in the Point Clouds: Studying the Robustness of Point Cloud Convolutions
- Integrating Multiple Receptive Fields through Grouped Active Convolution
- Tube-CNN: Modeling temporal evolution of appearance for object detection in video
- Social Behavioral Phenotyping of Drosophila with a2D-3D Hybrid CNN Framework
- CSST Slitless Spectra: Target Detection and Classification with YOLO
- Collaborative Deep Reinforcement Learning for Joint Object Search
- Convolutional Channel Features
- Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object Detection
- RRNet: Repetition-Reduction Network for Energy Efficient Decoder of Depth Estimation
- Bypassing the static input size of neural networks in flare forecasting by using spatial pyramid pooling
- CNNCat: Categorizing high-energy photons in a Compton/Pair Telescope with Convolutional Neural Networks
- Anysize GAN: A solution to the image-warping problem
- An active search strategy for efficient object class detection
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- Joint Facade Registration and Segmentation for Urban Localization
- Multi-Scale Aligned Distillation for Low-Resolution Detection
- DeepKey: Towards End-to-End Physical Key Replication From a Single Photograph
- Non-local RoI for Cross-Object Perception
- Multi-Scale Weight Sharing Network for Image Recognition
- A Video Analysis Method on Wanfang Dataset via Deep Neural Network
- A new smart-cropping pipeline for prostate segmentation using deep learning networks
- Multi-Scale Dual-Branch Fully Convolutional Network for Hand Parsing
- Quantized Neural Networks via {-1, +1} Encoding Decomposition and Acceleration
- The Compressed Model of Residual CNDS
- Weakly Supervised Instance Segmentation by Deep Community Learning
- Emergence of Selective Invariance in Hierarchical Feed Forward Networks
- Visual Search at Pinterest
- Harvesting Visual Objects from Internet Images via Deep Learning Based Objectness Assessment
- Matching-space Stereo Networks for Cross-domain Generalization
- Comparison of the Deep-Learning-Based Automated Segmentation Methods for the Head Sectioned Images of the Virtual Korean Human Project
- Improved Super-Resolution Convolution Neural Network for Large Images
- DropRegion Training of Inception Font Network for High-Performance Chinese Font Recognition
- LiDAR point-cloud processing based on projection methods: a comparison
- Residual-CNDS for Grand Challenge Scene Dataset
- A Hybrid Framework for Matching Printing Design Files to Product Photos
- Learning to Compose with Professional Photographs on the Web
- S-OHEM: Stratified Online Hard Example Mining for Object Detection
- Per-Pixel Feedback for improving Semantic Segmentation
- DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm Retrieval
- Deep Image Deraining Via Intrinsic Rainy Image Priors and Multi-scale Auxiliary Decoding
- Multi-Scale Spatially-Asymmetric Recalibration for Image Classification
- TS2C: Tight Box Mining with Surrounding Segmentation Context for Weakly Supervised Object Detection
- Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
- Learning to Segment Object Candidates via Recursive Neural Networks
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- Semantic Segmentation for Urban-Scene Images
- Region-based semantic segmentation with end-to-end training
- Context-LGM: Leveraging Object-Context Relation for Context-Aware Object Recognition