Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports
arXiv:2111.03452 · doi:10.1038/s42256-021-00425-9
Abstract
Pre-training lays the foundation for recent successes in radiograph analysis supported by deep learning. It learns transferable image representations by conducting large-scale fully-supervised or self-supervised learning on a source domain. However, supervised pre-training requires a complex and labor intensive two-stage human-assisted annotation process while self-supervised learning cannot compete with the supervised paradigm. To tackle these issues, we propose a cross-supervised methodology named REviewing FreE-text Reports for Supervision (REFERS), which acquires free supervision signals from original radiology reports accompanying the radiographs. The proposed approach employs a vision transformer and is designed to learn joint representations from multiple views within every patient study. REFERS outperforms its transfer learning and self-supervised learning counterparts on 4 well-known X-ray datasets under extremely limited supervision. Moreover, REFERS even surpasses methods based on a source domain of radiographs with human-assisted structured labels. Thus REFERS has the potential to replace canonical pre-training methodologies.
Accepted by Nature Machine Intelligence. The official version is at https://www.nature.com/articles/s42256-021-00425-9. Codes are available at https://github.com/funnyzhou/REFERS
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- How transferable are features in deep neural networks?
- cuDNN: Efficient Primitives for Deep Learning
- COVID-19 Image Data Collection: Prospective Predictions Are the Future
- Models Genesis
- VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations
- CheXphoto: 10,000+ Photos and Transformations of Chest X-rays for Benchmarking Deep Learning Robustness
Cited by in corpus (9)
- Is attention all you need in medical image analysis? A review
- CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training
- CLIP in Medical Imaging: A Survey
- Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning
- Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training
- Adaptive Betweenness Clustering for Semi-Supervised Domain Adaptation
- Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
- Enhancing medical vision-language contrastive learning via inter-matching relation modelling
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation