GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data
arXiv:2402.14973 · doi:10.1016/j.csl.2025.101785
Abstract
Multimodal Large Language Models (MLLMs) are typically assessed using expensive annotated multimodal benchmarks, which often lag behind the rapidly evolving demands of MLLM evaluation. This paper outlines and validates GenCeption, a novel, annotation-free evaluation method that requires only unimodal data to measure inter-modality semantic coherence and inversely assesses MLLMs' tendency to hallucinate. This approach eliminates the need for costly data annotation, minimizes the risk of training data contamination, is expected to result in slower benchmark saturation, and avoids the illusion of emerging abilities. Inspired by the DrawCeption game, GenCeption begins with a non-textual sample and proceeds through iterative description and generation steps. The semantic drift across iterations is quantified using the GC@T metric. While GenCeption is principally applicable to MLLMs across various modalities, this paper focuses on its implementation and validation for Vision LLMs (VLLMs). Based on the GenCeption method, we establish the MMECeption benchmark for evaluating VLLMs, and compare the performance of several popular VLLMs and human annotators. Our empirical results validate GenCeption's effectiveness, demonstrating strong correlations with established VLLM benchmarks. VLLMs still significantly lag behind human performance and struggle especially with text-intensive tasks.
Published by Computer Speech & Language (https://doi.org/10.1016/j.csl.2025.101785). Source code and Leaderboard: https://github.com/llcresearch/GenCeption
References in corpus (15)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Are Emergent Abilities of Large Language Models a Mirage?
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- On Evaluating Adversarial Robustness of Large Vision-Language Models
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
- Emu3: Next-Token Prediction is All You Need
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models