Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
arXiv:2508.19294 · doi:10.1016/j.inffus.2025.103575
Abstract
The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This in-depth review presents a structured exploration of the state-of-the-art in LVLMs, systematically organized through a three-step research review process. First, we discuss the functioning of vision language models (VLMs) for object detection, describing how these models harness natural language processing (NLP) and computer vision (CV) techniques to revolutionize object detection and localization. We then explain the architectural innovations, training paradigms, and output flexibility of recent LVLMs for object detection, highlighting how they achieve advanced contextual understanding for object detection. The review thoroughly examines the approaches used in integration of visual and textual information, demonstrating the progress made in object detection using VLMs that facilitate more sophisticated object detection and localization strategies. This review presents comprehensive visualizations demonstrating LVLMs' effectiveness in diverse scenarios including localization and segmentation, and then compares their real-time performance, adaptability, and complexity to traditional deep learning systems. Based on the review, its is expected that LVLMs will soon meet or surpass the performance of conventional methods in object detection. The review also identifies a few major limitations of the current LVLM modes, proposes solutions to address those challenges, and presents a clear roadmap for the future advancement in this field. We conclude, based on this study, that the recent advancement in LVLMs have made and will continue to make a transformative impact on object detection and robotic applications in the future.
First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)
References in corpus (31)
- LAION-5B: An open large-scale dataset for training next generation image-text models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- Toward Transformer-Based Object Detection
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- PaLI-3 Vision Language Models: Smaller, Faster, Stronger
- A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models
- Multimodal Foundation Models for Zero-shot Animal Species Recognition in Camera Trap Images
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
- Visual Large Language Models for Generalized and Specialized Applications
- Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
- LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
- OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
- Uncertainty-Aware Evaluation for Vision-Language Models
- VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
- Multimodal Autoregressive Pre-training of Large Vision Encoders
- REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
- LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
- Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
- A Simple Aerial Detection Baseline of Multimodal Language Models
- A Hitchhikers Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning
- Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
- AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks