YOLOX: Exceeding YOLO Series in 2021
arXiv:2107.08430
Abstract
In this report, we present some experienced improvements to YOLO series, forming a new high-performance detector -- YOLOX. We switch the YOLO detector to an anchor-free manner and conduct other advanced detection techniques, i.e., a decoupled head and the leading label assignment strategy SimOTA to achieve state-of-the-art results across a large scale range of models: For YOLO-Nano with only 0.91M parameters and 1.08G FLOPs, we get 25.3% AP on COCO, surpassing NanoDet by 1.8% AP; for YOLOv3, one of the most widely used detectors in industry, we boost it to 47.3% AP on COCO, outperforming the current best practice by 3.0% AP; for YOLOX-L with roughly the same amount of parameters as YOLOv4-CSP, YOLOv5-L, we achieve 50.0% AP on COCO at a speed of 68.9 FPS on Tesla V100, exceeding YOLOv5-L by 1.8% AP. Further, we won the 1st Place on Streaming Perception Challenge (Workshop on Autonomous Driving at CVPR 2021) using a single YOLOX-L model. We hope this report can provide useful experience for developers and researchers in practical scenes, and we also provide deploy versions with ONNX, TensorRT, NCNN, and Openvino supported. Source code is at https://github.com/Megvii-BaseDetection/YOLOX.
References in corpus (7)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Learning Spatial Fusion for Single-Shot Object Detection
- AutoAssign: Differentiable Label Assignment for Dense Object Detection
- Bag of Freebies for Training Object Detection Neural Networks
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- PP-YOLOv2: A Practical Object Detector
- LLA: Loss-aware Label Assignment for Dense Pedestrian Detection
Cited by in corpus (68)
- Towards Large-Scale Small Object Detection: Survey and Benchmarks
- Centralized Feature Pyramid for Object Detection
- CBNet: A Composite Backbone Network Architecture for Object Detection
- SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
- TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
- LXL: LiDAR Excluded Lean 3D Object Detection with 4D Imaging Radar and Camera Fusion
- A Survey on Deep Learning-Based Monocular Spacecraft Pose Estimation: Current State, Limitations and Prospects
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- Deep Learning in Automated Power Line Inspection: A Review
- Real-Time Accident Detection in Traffic Surveillance Using Deep Learning
- UnitModule: A Lightweight Joint Image Enhancement Module for Underwater Object Detection
- OGMN: Occlusion-guided Multi-task Network for Object Detection in UAV Images
- Continuous Human Action Recognition for Human-Machine Interaction: A Review
- Motion Robust High-Speed Light-Weighted Object Detection With Event Camera
- MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats
- Challenges for Monocular 6D Object Pose Estimation in Robotics
- iBall: Augmenting Basketball Videos with Gaze-moderated Embedded Visualizations
- Multitask Learning for SAR Ship Detection with Gaussian-Mask Joint Segmentation
- Vision meets algae: A novel way for microalgae recognization and health monitor
- Edge AI-Enabled Chicken Health Detection Based on Enhanced FCOS-Lite and Knowledge Distillation
- Task-wise Sampling Convolutions for Arbitrary-Oriented Object Detection in Aerial Images
- SoccerNet 2023 Challenges Results
- MPSN: Motion-aware Pseudo Siamese Network for Indoor Video Head Detection in Buildings
- Keypoint Promptable Re-Identification
- EARL: An Elliptical Distribution aided Adaptive Rotation Label Assignment for Oriented Object Detection in Remote Sensing Images
- Identification of Binary Neutron Star Mergers in Gravitational-Wave Data Using YOLO One-Shot Object Detection
- A Vision-Based Tactile Sensing System for Multimodal Contact Information Perception via Neural Network
- Contrastive Learning for Multi-Object Tracking with Transformers
- Learning Data Association for Multi-Object Tracking using Only Coordinates
- SEM-O-RAN: Semantic and Flexible O-RAN Slicing for NextG Edge-Assisted Mobile Systems
- A Multi-purpose Realistic Haze Benchmark with Quantifiable Haze Levels and Ground Truth
- Online V2X Scheduling for Raw-Level Cooperative Perception
- FCOSR: A Simple Anchor-free Rotated Detector for Aerial Object Detection
- Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
- Towards cumulative race time regression in sports: I3D ConvNet transfer learning in ultra-distance running events
- FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
- Towards pedestrian head tracking: A benchmark dataset and a multi-source data fusion network
- Table Detection for Visually Rich Document Images
- Optimal Kernel Orchestration for Tensor Programs with Korch
- Learnable Graph Matching: A Practical Paradigm for Data Association
- Tracking Skiers from the Top to the Bottom
- A Lightweight NMS-free Framework for Real-time Visual Fault Detection System of Freight Trains
- MuraNet: Multi-task Floor Plan Recognition with Relation Attention
- Adaptive Instance Distillation for Object Detection in Autonomous Driving
- Video object detection for privacy-preserving patient monitoring in intensive care
- CrowdSim2: an Open Synthetic Benchmark for Object Detectors
- YOLO11-JDE: Fast and Accurate Multi-Object Tracking with Self-Supervised Re-ID
- YOLIC: An Efficient Method for Object Localization and Classification on Edge Devices
- Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
- CAMOT: Camera Angle-aware Multi-Object Tracking
- Workshop on Autonomous Driving at CVPR 2021: Technical Report for Streaming Perception Challenge
- Monocular Per-Object Distance Estimation with Masked Object Modeling
- UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
- AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
- 1st Place Solution for the UVO Challenge on Image-based Open-World Segmentation 2021
- Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
- SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
- Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object Detection
- IDDR-NGP: Incorporating Detectors for Distractor Removal with Instant Neural Radiance Field
- To Make Yourself Invisible with Adversarial Semantic Contours
- Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
- Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety
- Automated Construction of Time-Space Diagrams for Traffic Analysis Using Street-View Video Sequence
- TrackID3x3: A Dataset and Algorithm for Multi-Player Tracking with Identification and Pose Estimation in 3x3 Basketball Full-court Videos
- A Multilevel Strategy to Improve People Tracking in a Real-World Scenario
- Towards Toxic and Narcotic Medication Detection with Rotated Object Detector
- UVO Challenge on Video-based Open-World Segmentation 2021: 1st Place Solution
- Segmentation-Based Bounding Box Generation for Omnidirectional Pedestrian Detection