MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis
arXiv:2506.08900 · doi:10.1038/s41746-025-01852-3
Abstract
Artificial intelligence (AI) has become a fundamental tool for assisting clinicians in analyzing ophthalmic images, such as optical coherence tomography (OCT). However, developing AI models often requires extensive annotation, and existing models tend to underperform on independent, unseen data. Foundation models (FMs), large AI models trained on vast unlabeled datasets, have shown promise in overcoming these challenges. Nonetheless, available FMs for ophthalmology lack extensive validation, especially for segmentation tasks, and focus on a single imaging modality. In this context, we propose MIRAGE, a novel multimodal FM for the analysis of OCT and scanning laser ophthalmoscopy (SLO) images. Additionally, we propose a new evaluation benchmark with OCT/SLO classification and segmentation tasks. The comparison with general and specialized FMs and segmentation methods shows the superiority of MIRAGE in both types of tasks, highlighting its suitability as a basis for the development of robust AI systems for retinal OCT image analysis. Both MIRAGE and the evaluation benchmark are publicly available: https://github.com/j-morano/MIRAGE.
Accepted for publication in npj Digital Medicine
References in corpus (13)
- Automated Design of Deep Learning Methods for Biomedical Image Segmentation
- Segment Anything in Medical Images
- Contrastive Representation Learning: A Framework and Review
- Segment Anything Model for Medical Image Analysis: an Experimental Study
- Segment Anything Model for Medical Images?
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- OCTDL: Optical Coherence Tomography Dataset for Image-Based Deep Learning Methods
- A Foundation Language-Image Model of the Retina (FLAIR): Encoding Expert Knowledge in Text Supervision
- VisionFM: a Multi-Modal Multi-Task Vision Foundation Model for Generalist Ophthalmic Artificial Intelligence
- Segmentation of Bruch's Membrane in retinal OCT with AMD using anatomical priors and uncertainty quantification
- RRWNet: Recursive Refinement Network for effective retinal artery/vein segmentation and classification
- Self-supervised learning via inter-modal reconstruction and feature projection networks for label-efficient 3D-to-2D segmentation
- Multimodal Transfer Learning-based Approaches for Retinal Vascular Segmentation