Contrastive Attention for Automatic Chest X-ray Report Generation
arXiv:2106.06965
Abstract
Recently, chest X-ray report generation, which aims to automatically generate descriptions of given chest X-ray images, has received growing research interests. The key challenge of chest X-ray report generation is to accurately capture and describe the abnormal regions. In most cases, the normal regions dominate the entire chest X-ray image, and the corresponding descriptions of these normal regions dominate the final report. Due to such data bias, learning-based models may fail to attend to abnormal regions. In this work, to effectively capture and describe abnormal regions, we propose the Contrastive Attention (CA) model. Instead of solely focusing on the current input image, the CA model compares the current input image with normal images to distill the contrastive information. The acquired contrastive information can better represent the visual features of abnormal regions. According to the experiments on the public IU-X-ray and MIMIC-CXR datasets, incorporating our CA into several existing models can boost their performance across most metrics. In addition, according to the analysis, the CA model can help existing models better attend to the abnormal regions and provide more accurate descriptions which are crucial for an interpretable diagnosis. Specifically, we achieve the state-of-the-art results on the two public datasets.
Appear in Findings of ACL 2021 (The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021))
References in corpus (10)
- Learning Transferable Visual Models From Natural Language Supervision
- Bootstrap your own latent: A new approach to self-supervised Learning
- Improved Baselines with Momentum Contrastive Learning
- Microsoft COCO Captions: Data Collection and Evaluation Server
- A Structured Self-attentive Sentence Embedding
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Contrastive Learning for Image Captioning
- Auxiliary Signal-Guided Knowledge Encoder-Decoder for Medical Report Generation
- Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation
- Addressing Data Bias Problems for Chest X-ray Image Report Generation