SSD: Single Shot MultiBox Detector
arXiv:1512.02325 · doi:10.1007/978-3-319-46448-0_2
Abstract
We present a method for detecting objects in images using a single deep neural network. Our approach, named SSD, discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per feature map location. At prediction time, the network generates scores for the presence of each object category in each default box and produces adjustments to the box to better match the object shape. Additionally, the network combines predictions from multiple feature maps with different resolutions to naturally handle objects of various sizes. Our SSD model is simple relative to methods that require object proposals because it completely eliminates proposal generation and subsequent pixel or feature resampling stage and encapsulates all computation in a single network. This makes SSD easy to train and straightforward to integrate into systems that require a detection component. Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified framework for both training and inference. Compared to other single stage methods, SSD has much better accuracy, even with a smaller input image size. For input, SSD achieves 72.1% mAP on VOC2007 test at 58 FPS on a Nvidia Titan X and for input, SSD achieves 75.1% mAP, outperforming a comparable state of the art Faster R-CNN model. Code is available at https://github.com/weiliu89/caffe/tree/ssd .
ECCV 2016
Cited by in corpus (333)
- A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
- A Survey of Deep Learning Techniques for Autonomous Driving
- Cell Detection with Star-convex Polygons
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Gliding vertex on the horizontal bounding box for multi-oriented object detection
- TextBoxes++: A Single-Shot Oriented Scene Text Detector
- Road Damage Detection Using Deep Neural Networks with Images Captured Through a Smartphone
- Deep Audio-Visual Speech Recognition
- PlantDoc: A Dataset for Visual Plant Disease Detection
- Slim-neck by GSConv: A lightweight-design for real-time detector architectures
- T-CNN: Tubelets with Convolutional Neural Networks for Object Detection from Videos
- Slicing Aided Hyper Inference and Fine-tuning for Small Object Detection
- Automatic Ship Detection of Remote Sensing Images from Google Earth in Complex Scenes Based on Multi-Scale Rotation Dense Feature Pyramid Networks
- 3D-CVF: Generating Joint Camera and LiDAR Features Using Cross-View Spatial Feature Fusion for 3D Object Detection
- A Survey on Instance Segmentation: State of the art
- Vision-based Robotic Grasping From Object Localization, Object Pose Estimation to Grasp Estimation for Parallel Grippers: A Review
- YOLOP: You Only Look Once for Panoptic Driving Perception
- The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
- Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark
- MMRotate: A Rotated Object Detection Benchmark using PyTorch
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Centralized Feature Pyramid for Object Detection
- Waste detection in Pomerania: non-profit project for detecting waste in environment
- HIT-UAV: A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection
- Loss Functions and Metrics in Deep Learning
- A Review of Object Detection Models based on Convolutional Neural Network
- YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series
- TransCrowd: weakly-supervised crowd counting with transformers
- Fracture Detection in Pediatric Wrist Trauma X-ray Images Using YOLOv8 Algorithm
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Hyperspectral Classification Based on Lightweight 3-D-CNN With Transfer Learning
- An Efficient and Layout-Independent Automatic License Plate Recognition System Based on the YOLO detector
- DeepIM: Deep Iterative Matching for 6D Pose Estimation
- Object Detection Under Rainy Conditions for Autonomous Vehicles: A Review of State-of-the-Art and Emerging Techniques
- Few-Example Object Detection with Model Communication
- An Attention-Fused Network for Semantic Segmentation of Very-High-Resolution Remote Sensing Imagery
- A Comparative Study of Fruit Detection and Counting Methods for Yield Mapping in Apple Orchards
- Focal Inverse Distance Transform Maps for Crowd Localization
- Bringing AI To Edge: From Deep Learning's Perspective
- Deep-Learning-Based Image Segmentation Integrated with Optical Microscopy for Automatically Searching for Two-Dimensional Materials
- RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- Scalable Image Coding for Humans and Machines
- From Handcrafted to Deep Features for Pedestrian Detection: A Survey
- A Survey on Long-Tailed Visual Recognition
- A review on deep learning techniques for 3D sensed data classification
- DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic
- BigDL: A Distributed Deep Learning Framework for Big Data
- Real-Time Fruit Recognition and Grasping Estimation for Autonomous Apple Harvesting
- Edge Deep Learning in Computer Vision and Medical Diagnostics: A Comprehensive Survey
- A comprehensive review of Binary Neural Network
- Rain rendering for evaluating and improving robustness to bad weather
- RGB-D Inertial Odometry for a Resource-Restricted Robot in Dynamic Environments
- A Survey on Approximate Edge AI for Energy Efficient Autonomous Driving Services
- Exploring Sequence Feature Alignment for Domain Adaptive Detection Transformers
- Zero-Shot Detection
- YOLO-TLA: An Efficient and Lightweight Small Object Detection Model based on YOLOv5
- PolarDet: A Fast, More Precise Detector for Rotated Target in Aerial Images
- Automatic Discovery and Geotagging of Objects from Street View Imagery
- A Survey on Collaborative DNN Inference for Edge Intelligence
- Review of data analysis in vision inspection of power lines with an in-depth discussion of deep learning technology
- Aerial Images Processing for Car Detection using Convolutional Neural Networks: Comparison between Faster R-CNN and YoloV3
- Segmentation of cell-level anomalies in electroluminescence images of photovoltaic modules
- Ellipse R-CNN: Learning to Infer Elliptical Object from Clustering and Occlusion
- High precision control and deep learning-based corn stand counting algorithms for agricultural robot
- A Single-Shot Arbitrarily-Shaped Text Detector based on Context Attended Multi-Task Learning
- Uncertainty Estimation in One-Stage Object Detection
- Advances and Applications of Computer Vision Techniques in Vehicle Trajectory Generation and Surrogate Traffic Safety Indicators
- Deep Learning Techniques for In-Crop Weed Identification: A Review
- A Deep Learning Based Automatic Defect Analysis Framework for In-situ TEM Ion Irradiations
- Convolutional Neural Networks for Image-based Corn Kernel Detection and Counting
- InsPLAD: A Dataset and Benchmark for Power Line Asset Inspection in UAV Images
- Enabling Pedestrian Safety using Computer Vision Techniques: A Case Study of the 2018 Uber Inc. Self-driving Car Crash
- Eye in the Sky: Drone-Based Object Tracking and 3D Localization
- Detection and Classification of Astronomical Targets with Deep Neural Networks in Wide Field Small Aperture Telescopes
- Visual-tactile Fusion for Transparent Object Grasping in Complex Backgrounds
- Automatic extraction of road intersection points from USGS historical map series using deep convolutional neural networks
- PillarGrid: Deep Learning-based Cooperative Perception for 3D Object Detection from Onboard-Roadside LiDAR
- ViGT: Proposal-free Video Grounding with Learnable Token in Transformer
- Deep Learning-Based Quantification of Pulmonary Hemosiderophages in Cytology Slides
- Generative adversarial network with object detector discriminator for enhanced defect detection on ultrasonic B-scans
- Multi-Target Multi-Camera Tracking of Vehicles using Metadata-Aided Re-ID and Trajectory-Based Camera Link Model
- Tomato Maturity Recognition with Convolutional Transformers
- A Systematic IoU-Related Method: Beyond Simplified Regression for Better Localization
- Real-Time Illegal Parking Detection System Based on Deep Learning
- Visual diagnosis of the Varroa destructor parasitic mite in honeybees using object detector techniques
- Swin-transformer-yolov5 For Real-time Wine Grape Bunch Detection
- Real-time Plant Health Assessment Via Implementing Cloud-based Scalable Transfer Learning On AWS DeepLens
- Cluster Pruning: An Efficient Filter Pruning Method for Edge AI Vision Applications
- MIDV-2020: A Comprehensive Benchmark Dataset for Identity Document Analysis
- Weakly Supervised Object Detection in Artworks
- Continuous Human Action Recognition for Human-Machine Interaction: A Review
- Industrial Scene Text Detection with Refined Feature-attentive Network
- CrossRoI: Cross-camera Region of Interest Optimization for Efficient Real Time Video Analytics at Scale
- Hybrid SNN-ANN: Energy-Efficient Classification and Object Detection for Event-Based Vision
- Automated building image extraction from 360° panoramas for postdisaster evaluation
- Background Learnable Cascade for Zero-Shot Object Detection
- Efficient Few-Shot Object Detection via Knowledge Inheritance
- AIParsing: Anchor-free Instance-level Human Parsing
- Registration-free Face-SSD: Single shot analysis of smiles, facial attributes, and affect in the wild
- LPF: A Language-Prior Feedback Objective Function for De-biased Visual Question Answering
- Enhancing Object Detection for Autonomous Driving by Optimizing Anchor Generation and Addressing Class Imbalance
- Oriented object detection in optical remote sensing images using deep learning: a survey
- MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats
- Traffic-Aware Multi-Camera Tracking of Vehicles Based on ReID and Camera Link Model
- Gradient-based Camera Exposure Control for Outdoor Mobile Platforms
- Online Ensemble Model Compression using Knowledge Distillation
- RS-YOLOX: A High Precision Detector for Object Detection in Satellite Remote Sensing Images
- DOLPHINS: Dataset for Collaborative Perception enabled Harmonious and Interconnected Self-driving
- Object condensation: one-stage grid-free multi-object reconstruction in physics detectors, graph and image data
- A Survey on Open Set Recognition
- Diverse Human Motion Prediction Guided by Multi-Level Spatial-Temporal Anchors
- Foreign-Object Detection in High-Voltage Transmission Line Based on Improved YOLOv8m
- YoTube: Searching Action Proposal via Recurrent and Static Regression Networks
- A Mobile App for Wound Localization using Deep Learning
- Detection of 3D Bounding Boxes of Vehicles Using Perspective Transformation for Accurate Speed Measurement
- Marvin: an Innovative Omni-Directional Robotic Assistant for Domestic Environments
- Detecting Solar-like Oscillations in Red Giants with Deep Learning
- CT-Net: Arbitrary-Shaped Text Detection via Contour Transformer
- Dynamic Label Assignment for Object Detection by Combining Predicted IoUs and Anchor IoUs
- Revisiting Computer-Aided Tuberculosis Diagnosis
- Towards Better Accuracy-efficiency Trade-offs: Divide and Co-training
- Panoptic Instance Segmentation on Pigs
- Extending Maps with Semantic and Contextual Object Information for Robot Navigation: a Learning-Based Framework using Visual and Depth Cues
- Vision meets algae: A novel way for microalgae recognization and health monitor
- Reliable Real Time Ball Tracking for Robot Table Tennis
- URBAN-i: From urban scenes to mapping slums, transport modes, and pedestrians in cities using deep learning and computer vision
- AirDet: Few-Shot Detection without Fine-tuning for Autonomous Exploration
- Split Computing for Complex Object Detectors: Challenges and Preliminary Results
- AM-MobileNet1D: A Portable Model for Speaker Recognition
- Rethinking Features-Fused-Pyramid-Neck for Object Detection
- Dual Semantic Fusion Network for Video Object Detection
- Precise Temporal Action Localization by Evolving Temporal Proposals
- Single-Shot Two-Pronged Detector with Rectified IoU Loss
- Comprehensive Performance Evaluation of YOLOv12, YOLO11, YOLOv10, YOLOv9 and YOLOv8 on Detecting and Counting Fruitlet in Complex Orchard Environments
- Unconstrained Face-Mask & Face-Hand Datasets: Building a Computer Vision System to Help Prevent the Transmission of COVID-19
- Learning to Detect Open Carry and Concealed Object with 77GHz Radar
- HANDS: A Multimodal Dataset for Modeling Towards Human Grasp Intent Inference in Prosthetic Hands
- Context in object detection: a systematic literature review
- High Performance Visual Tracking with Circular and Structural Operators
- E-UAV: An Edge-based Energy-Efficient Object Detection System for Unmanned Aerial Vehicles
- Reviewing Intelligent Cinematography: AI research for camera-based video production
- LGNN: A Context-aware Line Segment Detector
- Real-Time and Accurate Object Detection in Compressed Video by Long Short-term Feature Aggregation
- Learning from THEODORE: A Synthetic Omnidirectional Top-View Indoor Dataset for Deep Transfer Learning
- Learning Semantics-aware Distance Map with Semantics Layering Network for Amodal Instance Segmentation
- CSC-Unet: A Novel Convolutional Sparse Coding Strategy Based Neural Network for Semantic Segmentation
- Optical Transient Object Classification in Wide Field Small Aperture Telescopes with Neural Networks
- Towards Balanced Learning for Instance Recognition
- Vehicle Perception from Satellite
- FusionVision: A comprehensive approach of 3D object reconstruction and segmentation from RGB-D cameras using YOLO and fast segment anything
- Detecting Small Objects in Thermal Images Using Single-Shot Detector
- Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
- Blind-Spot Collision Detection System for Commercial Vehicles Using Multi Deep CNN Architecture
- Residual Pyramid Learning for Single-Shot Semantic Segmentation
- Active and Incremental Learning with Weak Supervision
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
- On the safety of vulnerable road users by cyclist orientation detection using Deep Learning
- Flexible and Fully Quantized Ultra-Lightweight TinyissimoYOLO for Ultra-Low-Power Edge Systems
- Radio astronomical images object detection and segmentation: A benchmark on deep learning methods
- Transformer-based Context Condensation for Boosting Feature Pyramids in Object Detection
- NaturalAE: Natural and Robust Physical Adversarial Examples for Object Detectors
- Versatile Multilinked Aerial Robot with Tilting Propellers: Design, Modeling, Control and State Estimation for Autonomous Flight and Manipulation
- CBNet: A Plug-and-Play Network for Segmentation-Based Scene Text Detection
- YOLOpeds: Efficient Real-Time Single-Shot Pedestrian Detection for Smart Camera Applications
- Deep Learning-Based Connector Detection for Robotized Assembly of Automotive Wire Harnesses
- Knowledge Amalgamation for Object Detection with Transformers
- Rethinking Cross-Domain Pedestrian Detection: A Background-Focused Distribution Alignment Framework for Instance-Free One-Stage Detectors
- Object Detector Differences when using Synthetic and Real Training Data
- Moderately Supervised Learning: Definition, Framework and Generality
- Group channel pruning and spatial attention distilling for object detection
- Self-Supervised Velocity Estimation for Automotive Radar Object Detection Networks
- Integrating GAN and Texture Synthesis for Enhanced Road Damage Detection
- A Specific Task-oriented Semantic Image Communication System for substation patrol inspection
- Orientation Aware Weapons Detection In Visual Data : A Benchmark Dataset
- Towards Accurate Human Pose Estimation in Videos of Crowded Scenes
- KOLOMVERSE: Korea open large-scale image dataset for object detection in the maritime universe
- Hand gesture detection in tests performed by older adults
- High-temporal-resolution event-based vehicle detection and tracking
- Viewpoint-driven Formation Control of Airships for Cooperative Target Tracking
- Visual Relationship Detection with Relative Location Mining
- Recognition of Multiple Food Items in a Single Photo for Use in a Buffet-Style Restaurant
- FBD-SV-2024: Flying Bird Object Detection Dataset in Surveillance Video
- Going beyond Free Viewpoint: Creating Animatable Volumetric Video of Human Performances
- BdSL36: A Dataset for Bangladeshi Sign Letters Recognition
- Transfer Learning for Instance Segmentation of Waste Bottles using Mask R-CNN Algorithm
- UAV-VLN: End-to-End Vision Language guided Navigation for UAVs
- Real-time Aerial Detection and Reasoning on Embedded-UAVs
- A Cost-Effective Person-Following System for Assistive Unmanned Vehicles with Deep Learning at the Edge
- Pack and Detect: Fast Object Detection in Videos Using Region-of-Interest Packing
- Optimisation of the PointPillars network for 3D object detection in point clouds
- Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability
- U-DECN: End-to-End Underwater Object Detection ConvNet with Improved DeNoising Training
- AD-Det: Boosting Object Detection in UAV Images with Focused Small Objects and Balanced Tail Classes
- A Robust Deep Networks based Multi-Object MultiCamera Tracking System for City Scale Traffic
- PIDNet: An Efficient Network for Dynamic Pedestrian Intrusion Detection
- MRZ code extraction from visa and passport documents using convolutional neural networks
- Potential Escalator-related Injury Identification and Prevention Based on Multi-module Integrated System for Public Health
- Face Detection in the Operating Room: Comparison of State-of-the-art Methods and a Self-supervised Approach
- Deep Learning Based Vehicle Make-Model Classification
- Detecting Anomalies in Software Execution Logs with Siamese Network
- Object Detection in Thermal Images Using Deep Learning for Unmanned Aerial Vehicles
- Assisting Blind People Using Object Detection with Vocal Feedback
- ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-based Systems
- Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level Annotations
- DSORT-MCU: Detecting Small Objects in Real-Time on Microcontroller Units
- aUToLights: A Robust Multi-Camera Traffic Light Detection and Tracking System
- Automotive Object Detection via Learning Sparse Events by Spiking Neurons
- LookOut! Interactive Camera Gimbal Controller for Filming Long Takes
- OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
- Prescriptive and Descriptive Approaches to Machine-Learning Transparency
- Advances in Deep Learning for Hyperspectral Image Analysis--Addressing Challenges Arising in Practical Imaging Scenarios
- DeepLight: Robust & Unobtrusive Real-time Screen-Camera Communication for Real-World Displays
- Learned Scalable Video Coding For Humans and Machines
- Autonomous Robotic Drilling System for Mice Cranial Window Creation: An Evaluation with an Egg Model
- DeePLT: Personalized Lighting Facilitates by Trajectory Prediction of Recognized Residents in the Smart Home
- Point Proposal Network for Reconstructing 3D Particle Endpoints with Sub-Pixel Precision in Liquid Argon Time Projection Chambers
- Towards Phytoplankton Parasite Detection Using Autoencoders
- PrivPAS: A real time Privacy-Preserving AI System and applied ethics
- Spatial-Temporal Deep Embedding for Vehicle Trajectory Reconstruction from High-Angle Video
- Large-Scale Classification of Structured Objects using a CRF with Deep Class Embedding
- Cut-and-Paste Dataset Generation for Balancing Domain Gaps in Object Instance Detection
- Visual inspection for illicit items in X-ray images using Deep Learning
- Wildfire Smoke Detection System: Model Architecture, Training Mechanism, and Dataset
- Tracking Skiers from the Top to the Bottom
- Synthetic Data-based Detection of Zebras in Drone Imagery
- SocialGuard: An Adversarial Example Based Privacy-Preserving Technique for Social Images
- Toward Accurate Person-level Action Recognition in Videos of Crowded Scenes
- Work-Efficient Parallel Non-Maximum Suppression Kernels
- Adaptive Instance Distillation for Object Detection in Autonomous Driving
- PLayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
- Six-channel Image Representation for Cross-domain Object Detection
- Real-Time AIoT for AAV Antenna Interference Detection via Edge-Cloud Collaboration
- Deep Multiple Instance Learning for Airplane Detection in High Resolution Imagery
- Efficient Perception, Planning, and Control Algorithm for Vision-Based Automated Vehicles
- Unmanned Aerial Vehicle Visual Detection and Tracking using Deep Neural Networks: A Performance Benchmark
- CAMOT: Camera Angle-aware Multi-Object Tracking
- hSDB-instrument: Instrument Localization Database for Laparoscopic and Robotic Surgeries
- TLD-READY: Traffic Light Detection -- Relevance Estimation and Deployment Analysis
- Mirror-Yolo: A Novel Attention Focus, Instance Segmentation and Mirror Detection Model
- Context-Aware Chart Element Detection
- Hybrid CNN Based Attention with Category Prior for User Image Behavior Modeling
- EditDuet: A Multi-Agent System for Video Non-Linear Editing
- Multiscale Detection of Cancerous Tissue in High Resolution Slide Scans
- MF-LPR: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
- Colonoscopy Polyp Detection and Classification: Dataset Creation and Comparative Evaluations
- Traffic Cameras to detect inland waterway barge traffic: An Application of machine learning
- FogGuard: guarding YOLO against fog using perceptual loss
- Learning Efficient Convolutional Networks through Irregular Convolutional Kernels
- Deep Instance Segmentation and Visual Servoing to Play Jenga with a Cost-Effective Robotic System
- Contrastive Learning and Cycle Consistency-based Transductive Transfer Learning for Target Annotation
- Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
- Detect-and-describe: Joint learning framework for detection and description of objects
- Frontiers in Intelligent Colonoscopy
- Neonatal Face and Facial Landmark Detection from Video Recordings
- Road Rutting Detection using Deep Learning on Images
- Infrared bubble recognition in the Milky Way and beyond using deep learning
- FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
- Development of a face mask detection pipeline for mask-wearing monitoring in the era of the COVID-19 pandemic: A modular approach
- An Efficient Deep Learning-Based Approach to Automating Invoice Document Validation
- Prediction Accuracy & Reliability: Classification and Object Localization under Distribution Shift
- Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
- Object Detection Based Handwriting Localization
- Accelerated Video Annotation driven by Deep Detector and Tracker
- A novel open-source ultrasound dataset with deep learning benchmarks for spinal cord injury localization and anatomical segmentation
- Illicit object detection in X-ray images using Vision Transformers
- Jointly Optimizing Sensing Pipelines for Multimodal Mixed Reality Interaction
- Real-time Embedded Person Detection and Tracking for Shopping Behaviour Analysis
- Jet Single Shot Detection
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
- THOR2: Topological Analysis for 3D Shape and Color-Based Human-Inspired Object Recognition in Unseen Environments
- Transforming a rare event search into a not-so-rare event search in real-time with deep learning-based object detection
- A Review and Implementation of Object Detection Models and Optimizations for Real-time Medical Mask Detection during the COVID-19 Pandemic
- Unit panel nodes detection by CNN on FAST reflector
- A Large Vision-Language Model based Environment Perception System for Visually Impaired People
- Bounding-box deep calibration for high performance face detection
- Development and Adaptation of Robotic Vision in the Real-World: the Challenge of Door Detection
- A Proper Orthogonal Decomposition approach for parameters reduction of Single Shot Detector networks
- A Simple Baseline for Pose Tracking in Videos of Crowded Scenes
- FashionFail: Addressing Failure Cases in Fashion Object Detection and Segmentation
- Less is More: Accelerating Faster Neural Networks Straight from JPEG
- SALINA: Towards Sustainable Live Sonar Analytics in Wild Ecosystems
- UESegNet: Context Aware Unconstrained ROI Segmentation Networks for Ear Biometric
- AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model
- Uncertainty-Aware Knowledge Distillation for Compact and Efficient 6DoF Pose Estimation
- Dual-Stream Spectral Decoupling Distillation for Remote Sensing Object Detection
- Topological Navigation Graph Framework
- The Application of Deep Learning for Lymph Node Segmentation: A Systematic Review
- FeDETR: a Federated Approach for Stenosis Detection in Coronary Angiography
- To Make Yourself Invisible with Adversarial Semantic Contours
- Real-Time Oil Leakage Detection on Aftermarket Motorcycle Damping System with Convolutional Neural Networks
- Dynamic Supervisor for Cross-dataset Object Detection
- Autonomous Curiosity for Real-Time Training Onboard Robotic Agents
- Non-iterative optimization of pseudo-labeling thresholds for training object detection models from multiple datasets
- DrawMon: A Distributed System for Detection of Atypical Sketch Content in Concurrent Pictionary Games
- Optimized Loss Functions for Object detection: A Case Study on Nighttime Vehicle Detection
- Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety
- Improving Long-Tailed Object Detection with Balanced Group Softmax and Metric Learning
- Data-free mixed-precision quantization using novel sensitivity metric
- Behavioral Learning of Dish Rinsing and Scrubbing based on Interruptive Direct Teaching Considering Assistance Rate
- Advancing biological super-resolution microscopy through deep learning: a brief review
- EC-IoU: Orienting Safety for Object Detectors via Ego-Centric Intersection-over-Union
- A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
- Overcoming the Limitations of Localization Uncertainty: Efficient & Exact Non-Linear Post-Processing and Calibration
- SPACE-SUIT: An Artificial Intelligence Based Chromospheric Feature Extractor and Classifier for SUIT
- Robust Human Identity Anonymization using Pose Estimation
- Detecting soccer balls with reduced neural networks: a comparison of multiple architectures under constrained hardware scenarios
- DynaMO: Protecting Mobile DL Models through Coupling Obfuscated DL Operators
- Access Control with Encrypted Feature Maps for Object Detection Models
- Semi-supervised Learning From Demonstration Through Program Synthesis: An Inspection Robot Case Study
- An Efficient and Generalizable Transfer Learning Method for Weather Condition Detection on Ground Terminals
- Instance-Level Safety-Aware Fidelity of Synthetic Data and Its Calibration
- ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks
- A Robust Deep Learning Framework for Prominence Detection through Composite Feature Representations
- CD-TWINSAFE: A ROS-enabled Digital Twin for Scene Understanding and Safety Emerging V2I Technology
- Detection of preventable fetal distress during labor from scanned cardiotocogram tracings using deep learning
- Intra-class Patch Swap for Self-Distillation
- Synthetic Data for Robust Runway Detection
- KAMERA: Enhancing Aerial Surveys of Ice-associated Seals in Arctic Environments
- Event-based Civil Infrastructure Visual Defect Detection: ev-CIVIL Dataset and Benchmark
- Enriching physical-virtual interaction in AR gaming by tracking identical objects via an egocentric partial observation frame
- Self-Guided Multiple Instance Learning for Weakly Supervised Disease Classification and Localization in Chest Radiographs
- MaskBEV: Joint Object Detection and Footprint Completion for Bird's-eye View 3D Point Clouds
- Assessing the Uncertainty and Robustness of the Laptop Refurbishing Software
- Web-Scale Generic Object Detection at Microsoft Bing
- Spot What Matters: Learning Context Using Graph Convolutional Networks for Weakly-Supervised Action Detection
- Sequence-SOD: Bio-inspired Sequence-aware Spiking ObjectDetection for Event Cameras
- Maritime Small Object Detection from UAVs using Deep Learning with Altitude-Aware Dynamic Tiling
- Open Vocabulary Word Recognition From Transcribed Bangla Texts
- A Comprehensive Framework for Automated Quality Control in the Automotive Industry
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
- A Cloud-Based Hybrid Model for Real-Time Detection of BRTA-Approved Licence Plates Using YOLO Tiny and Haar Cascade