BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
arXiv:2305.17100 · doi:10.1038/s41591-024-03185-2
Abstract
Traditional biomedical artificial intelligence (AI) models, designed for specific tasks or modalities, often exhibit limited flexibility in real-world deployment and struggle to utilize holistic information. Generalist AI holds the potential to address these limitations due to its versatility in interpreting different data types and generating tailored outputs for diverse needs. However, existing biomedical generalist AI solutions are typically heavyweight and closed source to researchers, practitioners, and patients. Here, we propose BiomedGPT, the first open-source and lightweight vision-language foundation model, designed as a generalist capable of performing various biomedical tasks. BiomedGPT achieved state-of-the-art results in 16 out of 25 experiments while maintaining a computing-friendly model scale. We also conducted human evaluations to assess the capabilities of BiomedGPT in radiology visual question answering, report generation, and summarization. BiomedGPT exhibits robust prediction ability with a low error rate of 3.8% in question answering, satisfactory performance with an error rate of 8.3% in writing complex radiology reports, and competitive summarization ability with a nearly equivalent preference score to human experts. Our method demonstrates that effective training with diverse data can lead to more practical biomedical AI for improving diagnosis and workflow efficiency.
Fix incorrect citations and add journal reference for the published version. Nat Med (2024)
References in corpus (20)
- LLaMA: Open and Efficient Foundation Language Models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classification
- Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
- MedViT: A Robust Vision Transformer for Generalized Medical Image Classification
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT
- Pix2seq: A Language Modeling Framework for Object Detection
- SciFive: a text-to-text transformer model for biomedical literature
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
- Multimodal Image-Text Matching Improves Retrieval-based Chest X-Ray Report Generation
- Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation
- A Lightweight, Rapid and Efficient Deep Convolutional Network for Chest X-Ray Tuberculosis Detection
- Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
- RadAdapt: Radiology Report Summarization via Lightweight Domain Adaptation of Large Language Models
- Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards
Cited by in corpus (12)
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- Artificial General Intelligence for Medical Imaging Analysis
- The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
- Improving Representation of High-frequency Components for Medical Visual Foundation Models
- Investigating the Quality of DermaMNIST and Fitzpatrick17k Dermatological Image Datasets
- A recent evaluation on the performance of LLMs on radiation oncology physics using questions of randomly shuffled options
- Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- Evidence Is All You Need: Ordering Imaging Studies via Language Model Alignment with the ACR Appropriateness Criteria
- Towards Cardiac MRI Foundation Models: Comprehensive Visual-Tabular Representations for Whole-Heart Assessment and Beyond
- A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
- Performance Assessment Strategies for Language Model Applications in Healthcare