Woodpecker: Hallucination Correction for Multimodal Large Language Models
arXiv:2310.16045 · doi:10.1007/s11432-024-4251-x
Abstract
Hallucination is a big shadow hanging over the rapidly evolving Multimodal Large Language Models (MLLMs), referring to the phenomenon that the generated text is inconsistent with the image content. In order to mitigate hallucinations, existing studies mainly resort to an instruction-tuning manner that requires retraining the models with specific data. In this paper, we pave a different way, introducing a training-free method named Woodpecker. Like a woodpecker heals trees, it picks out and corrects hallucinations from the generated text. Concretely, Woodpecker consists of five stages: key concept extraction, question formulation, visual knowledge validation, visual claim generation, and hallucination correction. Implemented in a post-remedy manner, Woodpecker can easily serve different MLLMs, while being interpretable by accessing intermediate outputs of the five stages. We evaluate Woodpecker both quantitatively and qualitatively and show the huge potential of this new paradigm. On the POPE benchmark, our method obtains a 30.66%/24.33% improvement in accuracy over the baseline MiniGPT-4/mPLUG-Owl. The source code is released at https://github.com/BradyFU/Woodpecker.
Accepted by Science China Information Sciences (SCIS)
References in corpus (35)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- A Survey on Multimodal Large Language Models
- Wizard of Wikipedia: Knowledge-Powered Conversational agents
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Improving language models by retrieving from trillions of tokens
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Perceiver: General Perception with Iterative Attention
- Fewer is More: Efficient Object Detection in Large Aerial Images
- BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
- Otter: A Multi-Modal Model with In-Context Instruction Tuning
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions
- SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
- Evaluation and Analysis of Hallucination in Large Vision-Language Models
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning
- Evaluating Object Hallucination in Large Vision-Language Models
- Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
- Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
- OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
- Detecting Everything in the Open World: Towards Universal Object Detection
- IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
- VIGC: Visual Instruction Generation and Correction
- Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
Cited by in corpus (5)
- A Survey on Multimodal Large Language Models
- A Large Vision-Language Model based Environment Perception System for Visually Impaired People
- MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
- Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
- VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing