Contrastive Learning of Medical Visual Representations from Paired Images and Text
arXiv:2010.00747
Abstract
Learning visual representations of medical images (e.g., X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights transferred from ImageNet pretraining, which is suboptimal due to drastically different image characteristics, or rule-based label extraction from the textual report data paired with medical images, which is inaccurate and hard to generalize. Meanwhile, several recent studies show exciting results from unsupervised contrastive learning from natural images, but we find these methods help little on medical images because of their high inter-class similarity. We propose ConVIRT, an alternative unsupervised strategy to learn medical visual representations by exploiting naturally occurring paired descriptive text. Our new method of pretraining medical image encoders with the paired text data via a bidirectional contrastive objective between the two modalities is domain-agnostic, and requires no additional expert input. We test ConVIRT by transferring our pretrained weights to 4 medical image classification tasks and 2 zero-shot retrieval tasks, and show that it leads to image representations that considerably outperform strong baselines in most settings. Notably, in all 4 classification tasks, our method requires only 10\% as much labeled training data as an ImageNet initialized counterpart to achieve better or comparable performance, demonstrating superior data efficiency.
First published in 2020. Accepted at Machine Learning for Healthcare (MLHC) 2022
References in corpus (1)
Cited by in corpus (42)
- Learning Transferable Visual Models From Natural Language Supervision
- Learning to Prompt for Vision-Language Models
- On the Opportunities and Risks of Foundation Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Recent advances and clinical applications of deep learning in medical image analysis
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Uncertainty-aware Contrastive Distillation for Incremental Semantic Segmentation
- Predicting post-operative right ventricular failure using video-based deep learning
- Augmenting Low-Resource Text Classification with Graph-Grounded Pre-training and Prompting
- A Review of Predictive and Contrastive Self-supervised Learning for Medical Images
- MERLOT: Multimodal Neural Script Knowledge Models
- Medical Imaging and Machine Learning
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- Contrastive Learning of Single-Cell Phenotypic Representations for Treatment Classification
- LAFITE: Towards Language-Free Training for Text-to-Image Generation
- RepsNet: Combining Vision with Language for Automated Medical Reports
- MedAug: Contrastive learning leveraging patient metadata improves representations for chest X-ray interpretation
- CLIP-ReIdent: Contrastive Training for Player Re-Identification
- Poisoning and Backdooring Contrastive Learning
- Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
- Clinically Labeled Contrastive Learning for OCT Biomarker Classification
- Text-to-Motion Retrieval: Towards Joint Understanding of Human Motion Data and Natural Language
- Self-supervised Image-text Pre-training With Mixed Data In Chest X-rays
- Automatic Medical Report Generation: Methods and Applications
- Imposing Relation Structure in Language-Model Embeddings Using Contrastive Learning
- Recent Advances in Medical Image Classification
- Pneumonia Detection on Chest X-ray using Radiomic Features and Contrastive Learning
- Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
- Modern Hopfield Networks for Few- and Zero-Shot Reaction Template Prediction
- Multimodal Representation Learning via Maximization of Local Mutual Information
- Deep AUC Maximization for Medical Image Classification: Challenges and Opportunities
- Semi-weakly Supervised Contrastive Representation Learning for Retinal Fundus Images
- Contrastive Video-Language Segmentation
- Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation
- CROCS: Clustering and Retrieval of Cardiac Signals Based on Patient Disease Class, Sex, and Age
- Unsupervised Local Discrimination for Medical Images
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
- Let Your Heart Speak in its Mother Tongue: Multilingual Captioning of Cardiac Signals
- An implementation of the "Guess who?" game using CLIP
- MIC: Model-agnostic Integrated Cross-channel Recommenders
- Overcoming the Domain Gap in Contrastive Learning of Neural Action Representations