Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
arXiv:2401.08396 · doi:10.1038/s41746-024-01185-7
Abstract
Recent studies indicate that Generative Pre-trained Transformer 4 with Vision (GPT-4V) outperforms human physicians in medical challenge tasks. However, these evaluations primarily focused on the accuracy of multi-choice questions alone. Our study extends the current scope by conducting a comprehensive analysis of GPT-4V's rationales of image comprehension, recall of medical knowledge, and step-by-step multimodal reasoning when solving New England Journal of Medicine (NEJM) Image Challenges - an imaging quiz designed to test the knowledge and diagnostic capabilities of medical professionals. Evaluation results confirmed that GPT-4V performs comparatively to human physicians regarding multi-choice accuracy (81.6% vs. 77.8%). GPT-4V also performs well in cases where physicians incorrectly answer, with over 78% accuracy. However, we discovered that GPT-4V frequently presents flawed rationales in cases where it makes the correct final choices (35.5%), most prominent in image comprehension (27.2%). Regardless of GPT-4V's high accuracy in multi-choice questions, our findings emphasize the necessity for further in-depth evaluations of its rationales before integrating such multimodal AI models into clinical workflows.
References in corpus (8)
- Capabilities of GPT-4 on Medical Challenge Problems
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- PubMed and Beyond: Biomedical Literature Search in the Age of Artificial Intelligence
- BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
- Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis
- Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V
- Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
Cited by in corpus (3)
- Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE Questions
- Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices for Generative AI
- Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images