Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing
arXiv:2204.09817 · doi:10.1007/978-3-031-20059-5_1
Abstract
Multi-modal data abounds in biomedicine, such as radiology images and reports. Interpreting this data at scale is essential for improving clinical care and accelerating clinical research. Biomedical text with its complex semantics poses additional challenges in vision--language modelling compared to the general domain, and previous work has used insufficiently adapted models that lack domain-specific language understanding. In this paper, we show that principled textual semantic modelling can substantially improve contrastive learning in self-supervised vision--language processing. We release a language model that achieves state-of-the-art results in radiology natural language inference through its improved vocabulary and novel language pretraining objective leveraging semantics and discourse characteristics in radiology reports. Further, we propose a self-supervised joint vision--language approach with a focus on better text modelling. It establishes new state of the art results on a wide range of publicly available benchmarks, in part by leveraging our new domain-specific language model. We release a new dataset with locally-aligned phrase grounding annotations by radiologists to facilitate the study of complex semantic modelling in biomedical vision--language processing. A broad evaluation, including on this new dataset, shows that our contrastive learning approach, aided by textual-semantic modelling, outperforms prior methods in segmentation tasks, despite only using a global-alignment objective.
To appear in ECCV 2022. Code: https://aka.ms/biovil-code Dataset: https://aka.ms/ms-cxr Demo Notebook: https://aka.ms/biovil-demo-notebook
References in corpus (1)
Cited by in corpus (20)
- CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training
- Automated Radiology Report Generation: A Review of Recent Advances
- Exploring scalable medical image encoders beyond text supervision
- A scoping review on multimodal deep learning in biomedical images and texts
- CAMANet: Class Activation Map Guided Attention Network for Radiology Report Generation
- Enhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive Learning
- Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
- PadChest-GR: A Bilingual Chest X-ray Dataset for Grounded Radiology Report Generation
- Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
- M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
- Enhancing medical vision-language contrastive learning via inter-matching relation modelling
- Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
- Radiomics-guided Multimodal Self-attention Network for Predicting Pathological Complete Response in Breast MRI
- Meta-Transfer Derm-Diagnosis: Exploring Few-Shot Learning and Transfer Learning for Skin Disease Classification in Long-Tail Distribution
- Using Multiple Instance Learning to Build Multimodal Representations
- Automated Spinal MRI Labelling from Reports Using a Large Language Model
- Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
- VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback