56 citations · 120 across the 15 of their papers we have counts for
9 papers · 1 filter
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
Ranjan Sapkota, Manoj Karkee
The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, a…
Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks
Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee +2
Automation in agriculture plays a vital role in addressing challenges related to crop monitoring and disease management, particularly through early detection systems. This study in…
A Review of 3D Object Detection with Vision-Language Models
Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally +2
This review provides a systematic analysis of comprehensive survey of 3D object detection with vision-language models(VLMs) , a rapidly advancing area at the intersection of 3D vis…
RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity
Ranjan Sapkota, Rahul Harsha Cheppally, Ajay Sharda +1
This study conducts a detailed comparison of RF-DETR object detection base model and YOLOv12 object detection model configurations for detecting greenfruits in a complex orchard en…
Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10
Ranjan Sapkota, Manoj Karkee
This study evaluated the performance of the YOLOv12 object detection model, and compared against the performances YOLOv11 and YOLOv10 for apple detection in commercial orchards bas…
Integrating YOLO11 and Convolution Block Attention Module for Multi-Season Segmentation of Tree Trunks and Branches in Commercial Apple Orchards
Ranjan Sapkota, Manoj Karkee
In this study, we developed a customized instance segmentation model by integrating the Convolutional Block Attention Module (CBAM) with the YOLO11 architecture. This model, traine…