VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback
arXiv:2501.17726 · doi:10.1016/j.mlwa.2025.100684
Abstract
As artificial intelligence (AI) becomes increasingly central to healthcare, the demand for explainable and trustworthy models is paramount. Current report generation systems for chest X-rays (CXR) often lack mechanisms for validating outputs without expert oversight, raising concerns about reliability and interpretability. To address these challenges, we propose a novel multimodal framework designed to enhance the semantic alignment and localization accuracy of AI-generated medical reports. Our framework integrates two key modules: a Phrase Grounding Model, which identifies and localizes pathologies in CXR images based on textual prompts, and a Text-to-Image Diffusion Module, which generates synthetic CXR images from prompts while preserving anatomical fidelity. By comparing features between the original and generated images, we introduce a dual-scoring system: one score quantifies localization accuracy, while the other evaluates semantic consistency. This approach significantly outperforms existing methods, achieving state-of-the-art results in pathology localization and text-to-image alignment. The integration of phrase grounding with diffusion models, coupled with the dual-scoring evaluation system, provides a robust mechanism for validating report quality, paving the way for more trustworthy and transparent AI in medical imaging.
References in corpus (8)
- Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for Radiology
- RoentGen: Vision-Language Foundation Model for Chest X-ray Generation
- Chest ImaGenome Dataset for Clinical Reasoning
- MAIRA-2: Grounded Radiology Report Generation
- XReal: Realistic Anatomy and Pathology-Aware X-ray Generation via Controllable Diffusion Model
- Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding
- How far generated data can impact Neural Networks performance?