Improving Chest X-Ray Report Generation by Leveraging Warm Starting
arXiv:2201.09405 · doi:10.1016/j.artmed.2023.102633
Abstract
Automatically generating a report from a patient's Chest X-Rays (CXRs) is a promising solution to reducing clinical workload and improving patient care. However, current CXR report generators -- which are predominantly encoder-to-decoder models -- lack the diagnostic accuracy to be deployed in a clinical setting. To improve CXR report generation, we investigate warm starting the encoder and decoder with recent open-source computer vision and natural language processing checkpoints, such as the Vision Transformer (ViT) and PubMedBERT. To this end, each checkpoint is evaluated on the MIMIC-CXR and IU X-Ray datasets. Our experimental investigation demonstrates that the Convolutional vision Transformer (CvT) ImageNet-21K and the Distilled Generative Pre-trained Transformer 2 (DistilGPT2) checkpoints are best for warm starting the encoder and decoder, respectively. Compared to the state-of-the-art ( Transformer Progressive), CvT2DistilGPT2 attained an improvement of 8.3\% for CE F-1, 1.8\% for BLEU-4, 1.6\% for ROUGE-L, and 1.0\% for METEOR. The reports generated by CvT2DistilGPT2 have a higher similarity to radiologist reports than previous approaches. This indicates that leveraging warm starting improves CXR report generation. Code and checkpoints for CvT2DistilGPT2 are available at https://github.com/aehrc/cvt2distilgpt2.
References in corpus (7)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Microsoft COCO Captions: Data Collection and Evaluation Server
- BEiT: BERT Pre-Training of Image Transformers
- Early Convolutions Help Transformers See Better
- Intriguing Properties of Vision Transformers
- XCiT: Cross-Covariance Image Transformers
- A Novel Collaborative Self-Supervised Learning Method for Radiomic Data
Cited by in corpus (10)
- Interactive and Explainable Region-guided Radiology Report Generation
- Automated Radiology Report Generation: A Review of Recent Advances
- Towards a Holistic Framework for Multimodal Large Language Models in Three-dimensional Brain CT Report Generation
- Automatic Medical Report Generation: Methods and Applications
- M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
- Structural Entities Extraction and Patient Indications Incorporation for Chest X-ray Report Generation
- Factual Serialization Enhancement: A Key Innovation for Chest X-ray Report Generation
- Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
- GIT-CXR: End-to-End Transformer for Chest X-Ray Report Generation
- Structure Observation Driven Image-Text Contrastive Learning for Computed Tomography Report Generation