A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
arXiv:2304.00501 · doi:10.3390/make5040083
Abstract
YOLO has become a central real-time object detection system for robotics, driverless cars, and video monitoring applications. We present a comprehensive analysis of YOLO's evolution, examining the innovations and contributions in each iteration from the original YOLO up to YOLOv8, YOLO-NAS, and YOLO with Transformers. We start by describing the standard metrics and postprocessing; then, we discuss the major changes in network architecture and training tricks for each model. Finally, we summarize the essential lessons from YOLO's development and provide a perspective on its future, highlighting potential research directions to enhance real-time object detection systems.
36 pages, 21 figures, 4 tables, published in Machine Learning and Knowledge Extraction. This version contains the last changes made to the published version
References in corpus (3)
Cited by in corpus (29)
- YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series
- Surround-View Fisheye Optics in Computer Vision and Simulation: Survey and Challenges
- On the Black-box Explainability of Object Detection Models for Safe and Trustworthy Industrial Applications
- Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving
- Efficient Vision-based Vehicle Speed Estimation
- Experimental validation of UAV search and detection system in real wilderness environment
- PerfCam: Digital Twinning for Production Lines Using 3D Gaussian Splatting and Vision Models
- NuSegDG: Integration of Heterogeneous Space and Gaussian Kernel for Domain-Generalized Nuclei Segmentation
- An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
- CCSPNet-Joint: Efficient Joint Training Method for Traffic Sign Detection Under Extreme Conditions
- Improving Small Drone Detection Through Multi-Scale Processing and Data Augmentation
- Real-Time Detection of Electronic Components in Waste Printed Circuit Boards: A Transformer-Based Approach
- An AI-Driven Multimodal Smart Home Platform for Continuous Monitoring and Assistance in Post-Stroke Motor Impairment
- Spatiotemporal Object Detection for Improved Aerial Vehicle Detection in Traffic Monitoring
- PILAR: Personalizing Augmented Reality Interactions with LLM-based Human-Centric and Trustworthy Explanations for Daily Use Cases
- TACO: Adversarial Camouflage Optimization on Trucks to Fool Object Detectors
- Efficient License Plate Recognition via Pseudo-Labeled Supervision with Grounding DINO and YOLOv8
- Layout-Aware OCR for Black Digital Archives with Unsupervised Evaluation
- IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
- Offloading Artificial Intelligence Workloads across the Computing Continuum by means of Active Storage Systems
- Vision-Language Assistant for Emotional Reactions to Risky Driving
- Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems
- SPACE-SUIT: An Artificial Intelligence Based Chromospheric Feature Extractor and Classifier for SUIT
- CSST Slitless Spectra: Target Detection and Classification with YOLO
- Improving Token-based Object Detection with Video
- Optimizing Helmet Detection with Hybrid YOLO Pipelines: A Detailed Analysis
- vNV-Heap: An Ownership-Based Virtually Non-Volatile Heap for Embedded Systems
- A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
- HitoMi-Cam: A Shape-Agnostic Person Detection Method Using the Spectral Characteristics of Clothing