ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs
arXiv:2305.15964 · doi:10.1109/TMI.2024.3398350
Abstract
The integration of Computer-Aided Diagnosis (CAD) with Large Language Models (LLMs) presents a promising frontier in clinical applications, notably in automating diagnostic processes akin to those performed by radiologists and providing consultations similar to a virtual family doctor. Despite the promising potential of this integration, current works face at least two limitations: (1) From the perspective of a radiologist, existing studies typically have a restricted scope of applicable imaging domains, failing to meet the diagnostic needs of different patients. Also, the insufficient diagnostic capability of LLMs further undermine the quality and reliability of the generated medical reports. (2) Current LLMs lack the requisite depth in medical expertise, rendering them less effective as virtual family doctors due to the potential unreliability of the advice provided during patient consultations. To address these limitations, we introduce ChatCAD+, to be universal and reliable. Specifically, it is featured by two main modules: (1) Reliable Report Generation and (2) Reliable Interaction. The Reliable Report Generation module is capable of interpreting medical images from diverse domains and generate high-quality medical reports via our proposed hierarchical in-context learning. Concurrently, the interaction module leverages up-to-date information from reputable medical websites to provide reliable medical advice. Together, these designed modules synergize to closely align with the expertise of human medical professionals, offering enhanced consistency and reliability for interpretation and advice. The source code is available at https://github.com/zhaozh10/ChatCAD.
Authors Zihao Zhao, Sheng Wang, Jinchen Gu, Yitao Zhu contributed equally to this work and should be considered co-first authors
References in corpus (14)
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- COVID-CT-Dataset: A CT Scan Dataset about COVID-19
- Skin Lesion Analysis toward Melanoma Detection: A Challenge at the International Symposium on Biomedical Imaging (ISBI) 2016, hosted by the International Skin Imaging Collaboration (ISIC)
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
- Follow My Eye: Using Gaze to Supervise Computer-Aided Diagnosis
- DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- Weakly Supervised Lesion Localization With Probabilistic-CAM Pooling
Cited by in corpus (5)
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- CLIP in Medical Imaging: A Survey
- MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models
- Analysing Environmental Efficiency in AI for X-Ray Diagnosis