1 citations · 1 across the 3 of their papers we have counts for
5 papers · 1 filter
SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation
Hyunwoong Kim, Seongeun Lee, Hannah Yun +2
Group Relative Policy Optimization (GRPO) and its variants, originally developed for Large Language Models (LLMs), have recently been applied to Multimodal LLMs and produced strong…
GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations
Jonggwon Park, Seongeun Lee, Junhyun Park +6
Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing rev…
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
Jonggwon Park, Byungmu Yoon, Soobum Kim +1
Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large language models (MLLMs) have enabled mul…
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
Jonggwon Park, Byungmu Yoon, Soobum Kim +1
Recent advancements in multimodal models have significantly improved vision-language (VL) alignment in radiology. However, existing approaches struggle to effectively utilize compl…
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
Jonggwon Park, Soobum Kim, Byungmu Yoon +2
The rapid evolution of artificial intelligence, especially in large language models (LLMs), has significantly impacted various domains, including healthcare. In chest X-ray (CXR) a…