Effectiveness Assessment of Recent Large Vision-Language Models
arXiv:2403.04306 · doi:10.1007/s44267-024-00050-1
Abstract
The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both specialized and general tasks warrants further investigation. This paper endeavors to evaluate the competency of popular LVLMs in specialized and general tasks, respectively, aiming to offer a comprehensive understanding of these novel models. To gauge their effectiveness in specialized tasks, we employ six challenging tasks in three different application scenarios: natural, healthcare, and industrial. These six tasks include salient/camouflaged/transparent object detection, as well as polyp detection, skin lesion detection, and industrial anomaly detection. We examine the performance of three recent open-source LVLMs, including MiniGPT-v2, LLaVA-1.5, and Shikra, on both visual recognition and localization in these tasks. Moreover, we conduct empirical investigations utilizing the aforementioned LVLMs together with GPT-4V, assessing their multi-modal understanding capabilities in general tasks including object counting, absurd question answering, affordance reasoning, attribute recognition, and spatial relation reasoning. Our investigations reveal that these LVLMs demonstrate limited proficiency not only in specialized tasks but also in general tasks. We delve deep into this inadequacy and uncover several potential factors, including limited cognition in specialized tasks, object hallucination, text-to-image interference, and decreased robustness in complex problems. We hope that this study can provide useful insights for the future development of LVLMs, helping researchers improve LVLMs for both general and specialized applications.
Accepted by Visual Intelligence
References in corpus (17)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Concealed Object Detection
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
- Structure-measure: A New Way to Evaluate Foreground Maps
- Light Field Salient Object Detection: A Review and Benchmark
- Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
- Salient Objects in Clutter: Bringing Salient Object Detection to the Foreground
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object Detection
- Siamese Network for RGB-D Salient Object Detection and Beyond
- How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges
- SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
- MOSE: A New Dataset for Video Object Segmentation in Complex Scenes
- Exposing and Mitigating Spurious Correlations for Cross-Modal Retrieval
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions