A comprehensive and easy-to-use multi-domain multi-task medical imaging meta-dataset
arXiv:2404.16000 · doi:10.1038/s41597-025-04866-4
Abstract
While the field of medical image analysis has undergone a transformative shift with the integration of machine learning techniques, the main challenge of these techniques is often the scarcity of large, diverse, and well-annotated datasets. Medical images vary in format, size, and other parameters and therefore require extensive preprocessing and standardization, for usage in machine learning. Addressing these challenges, we introduce the Medical Imaging Meta-Dataset (MedIMeta), a novel multi-domain, multi-task meta-dataset. MedIMeta contains 19 medical imaging datasets spanning 10 different domains and encompassing 54 distinct medical tasks, all of which are standardized to the same format and readily usable in PyTorch or other ML frameworks. We perform a technical validation of MedIMeta, demonstrating its utility through fully supervised and cross-domain few-shot learning baselines.
References in corpus (8)
- Overcoming catastrophic forgetting in neural networks
- The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
- The Liver Tumor Segmentation Benchmark (LiTS)
- MedMNIST v2 -- A large-scale lightweight benchmark for 2D and 3D biomedical image classification
- Meta-Learning for Semi-Supervised Few-Shot Classification
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Comparing Transfer and Meta Learning Approaches on a Unified Few-Shot Classification Benchmark
- Meta-Album: Multi-domain Meta-Dataset for Few-Shot Image Classification