Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
arXiv:2506.10825 · doi:10.1016/j.inffus.2025.103709
Abstract
Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer vision. The introduction of Segment Anything Model (SAM) set a milestone on segmentation of natural images, inspiring the design of a multitude of architectures for medical image segmentation. In this survey we offer a comprehensive and in-depth investigation on generalist models for medical image segmentation. We start with an introduction on the fundamentals concepts underpinning their development. Then, we provide a taxonomy on the different declinations of SAM in terms of zero-shot, few-shot, fine-tuning, adapters, on the recent SAM 2, on other innovative models trained on images alone, and others trained on both text and images. We thoroughly analyze their performances at the level of both primary research and best-in-literature, followed by a rigorous comparison with the state-of-the-art task-specific models. We emphasize the need to address challenges in terms of compliance with regulatory frameworks, privacy and security laws, budget, and trustworthy artificial intelligence (AI). Finally, we share our perspective on future directions concerning synthetic data, early fusion, lessons learnt from generalist models in natural language processing, agentic AI and physical AI, and clinical translation.
132 pages, 26 figures, 23 tables. Andrea Moglia and Matteo Leccardi are equally contributing authors
References in corpus (15)
- LoRA: Low-Rank Adaptation of Large Language Models
- Med3D: Transfer Learning for 3D Medical Image Analysis
- STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training
- A Comprehensive Survey on Segment Anything Model for Vision and Beyond
- Foundation Models for Biomedical Image Segmentation: A Survey
- s1: Simple test-time scaling
- Segment Anything in Medical Images and Videos: Benchmark and Deployment
- Universal Segmentation of 33 Anatomies
- Segment anything model 2: an application to 2D and 3D medical images
- Disruptive Autoencoders: Leveraging Low-level features for 3D Medical Image Pre-training
- Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
- How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
- Unleashing the Potential of SAM2 for Biomedical Images and Videos: A Survey
- Large-Scale 3D Medical Image Pre-training with Geometric Context Priors
- Vision Foundation Models in Medical Image Analysis: Advances and Challenges