Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
arXiv:2403.02469 · doi:10.3389/frai.2024.1430984
Abstract
Medical vision-language models (VLMs) combine computer vision (CV) and natural language processing (NLP) to analyze visual and textual medical data. Our paper reviews recent advancements in developing VLMs specialized for healthcare, focusing on models designed for medical report generation and visual question answering (VQA). We provide background on NLP and CV, explaining how techniques from both fields are integrated into VLMs to enable learning from multimodal data. Key areas we address include the exploration of medical vision-language datasets, in-depth analyses of architectures and pre-training strategies employed in recent noteworthy medical VLMs, and comprehensive discussion on evaluation metrics for assessing VLMs' performance in medical report generation and VQA. We also highlight current challenges and propose future directions, including enhancing clinical validity and addressing patient privacy concerns. Overall, our review summarizes recent progress in developing VLMs to harness multimodal medical data for improved healthcare applications.
43 pages; paper edited and restructured
References in corpus (62)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Training language models to follow instructions with human feedback
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- On the Opportunities and Risks of Foundation Models
- PaLM: Scaling Language Modeling with Pathways
- Flamingo: a Visual Language Model for Few-Shot Learning
- SurfaceNet: Adversarial SVBRDF Estimation from a Single Image
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Visual Instruction Tuning
- MedViT: A Robust Vision Transformer for Generalized Medical Image Classification
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Mistral 7B
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- VLP: A Survey on Vision-Language Pre-training
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training
- GIT: A Generative Image-to-text Transformer for Vision and Language
- Recurrent Neural Networks (RNNs): A gentle Introduction and Overview
- Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- A Comprehensive Survey of Continual Learning: Theory, Method and Application
- REFLACX, a dataset of reports and eye-tracking data for localization of abnormalities in chest x-rays
- Med-Flamingo: a Multimodal Medical Few-shot Learner
- A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
- XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
- MedCLIP: Contrastive Learning from Unpaired Medical Images and Text
- RepsNet: Combining Vision with Language for Automated Medical Reports
- Multimodal Image-Text Matching Improves Retrieval-based Chest X-Ray Report Generation
- A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics
- Masked Vision and Language Modeling for Multi-modal Representation Learning
- A Survey on Hallucination in Large Vision-Language Models
- Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models
- Investigating the Catastrophic Forgetting in Multimodal Large Language Models
- PMC-CLIP: Contrastive Language-Image Pre-training using Biomedical Documents
- Vision-Language Pre-training: Basics, Recent Advances, and Future Trends
- Improving Radiology Report Generation Systems by Removing Hallucinated References to Non-existent Priors
- Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing
- Aligning Large Multimodal Models with Factually Augmented RLHF
- RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
- Medical Vision Language Pretraining: A survey
- DePlot: One-shot visual language reasoning by plot-to-table translation
- Vision-Language Generative Model for View-Specific Chest X-ray Generation
- RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training
- A Systematic Review of Deep Learning-based Research on Radiology Report Generation
- Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT
- HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models
- Retrieving Multimodal Information for Augmented Generation: A Survey
- Embedding-based Multimodal Learning on Pan-Squamous Cell Carcinomas for Improved Survival Outcomes
- Privacy Preserving Federated Learning in Medical Imaging with Uncertainty Estimation
- Future-Proofing Medical Imaging with Privacy-Preserving Federated Learning and Uncertainty Quantification: A Review
- SA-Attack: Improving Adversarial Transferability of Vision-Language Pre-training Models via Self-Augmentation
- Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
- Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks
- Learning or Self-aligning? Rethinking Instruction Fine-tuning
- Brain-Inspired Continual Learning-Robust Feature Distillation and Re-Consolidation for Class Incremental Learning
- Secure Neuroimaging Analysis using Federated Learning with Homomorphic Encryption