Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
arXiv:2401.11421 · doi:10.1016/j.media.2024.103299
Abstract
Recently, vision-language representation learning has made remarkable advancements in building up medical foundation models, holding immense potential for transforming the landscape of clinical research and medical care. The underlying hypothesis is that the rich knowledge embedded in radiology reports can effectively assist and guide the learning process, reducing the need for additional labels. However, these reports tend to be complex and sometimes even consist of redundant descriptions that make the representation learning too challenging to capture the key semantic information. This paper develops a novel iterative vision-language representation learning framework by proposing a key semantic knowledge-emphasized report refinement method. Particularly, raw radiology reports are refined to highlight the key information according to a constructed clinical dictionary and two model-optimized knowledge-enhancement metrics. The iterative framework is designed to progressively learn, starting from gaining a general understanding of the patient's condition based on raw reports and gradually refines and extracts critical information essential to the fine-grained analysis tasks. The effectiveness of the proposed framework is validated on various downstream medical image analysis tasks, including disease classification, region-of-interest segmentation, and phrase grounding. Our framework surpasses seven state-of-the-art methods in both fine-tuning and zero-shot settings, demonstrating its encouraging potential for different clinical applications.
References in corpus (14)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Contrastive Learning of Medical Visual Representations from Paired Images and Text
- Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing
- MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports
- A coarse-to-fine framework for unsupervised multi-contrast MR image deformable registration with dual consistency constraint
- PyMIC: A deep learning toolkit for annotation-efficient medical image segmentation
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
- Masked Vision and Language Modeling for Multi-modal Representation Learning
- Advancing Radiograph Representation Learning with Masked Record Modeling