CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection
arXiv:2301.00785 · doi:10.1109/ICCV51070.2023.01934
Abstract
An increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited to segmenting specific organs/tumors and ignore the semantics of anatomical structures, nor can they be extended to novel domains. To address these issues, we propose the CLIP-Driven Universal Model, which incorporates text embedding learned from Contrastive Language-Image Pre-training (CLIP) to segmentation models. This CLIP-based label encoding captures anatomical relationships, enabling the model to learn a structured feature embedding and segment 25 organs and 6 types of tumors. The proposed model is developed from an assembly of 14 datasets, using a total of 3,410 CT scans for training and then evaluated on 6,162 external CT scans from 3 additional datasets. We rank first on the Medical Segmentation Decathlon (MSD) public leaderboard and achieve state-of-the-art results on Beyond The Cranial Vault (BTCV). Additionally, the Universal Model is computationally more efficient (6x faster) compared with dataset-specific models, generalized better to CT scans from varying sites, and shows stronger transfer learning performance on novel tasks.
ICCV-2023; Rank first in Medical Segmentation Decathlon (MSD) Competition
References in corpus (14)
- TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
- Cross-lingual Language Model Pretraining
- TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
- Med3D: Transfer Learning for 3D Medical Image Analysis
- AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation
- STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training
- Adapting Pretrained Vision-Language Foundational Models to Medical Imaging Domains
- MultiTalent: A Multi-Dataset Approach to Medical Image Segmentation
- Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?
- Medical Image Understanding with Pretrained Vision Language Models: A Comprehensive Study
- Towards General Purpose Medical AI: Continual Learning Medical Foundation Model
- Universal Lesion Detection by Learning from Multiple Heterogeneously Labeled Datasets
- Universal Segmentation of 33 Anatomies
- An End-to-End Framework For Universal Lesion Detection With Missing Annotations
Cited by in corpus (14)
- MONAI Label: A framework for AI-assisted Interactive Labeling of 3D Medical Images
- A Foundation Language-Image Model of the Retina (FLAIR): Encoding Expert Knowledge in Text Supervision
- CLIP in Medical Imaging: A Survey
- HybridMIM: A Hybrid Masked Image Modeling Framework for 3D Medical Image Segmentation
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- MultiTalent: A Multi-Dataset Approach to Medical Image Segmentation
- Data-Centric Foundation Models in Computational Healthcare: A Survey
- Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction
- Multi-organ segmentation: a progressive exploration of learning paradigms under scarce annotation
- Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation
- Acquiring Weak Annotations for Tumor Localization in Temporal and Volumetric Data
- AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
- Exploring the Transferability of a Foundation Model for Fundus Images: Application to Hypertensive Retinopathy
- Unified Medical Image Segmentation with State Space Modeling Snake