4 papers
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
Jifeng Song, Arun Das, Pan Wang +3
Scientific compound figures combine multiple labeled panels into a single image, and downstream pretraining and retrieval require panel-aligned visual-text pairs. However, in a Pub…
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
Kun Zhao, Guodong Liu, Hui Ji +6
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging du…
Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
Libin Qiu, Yuhang Ye, Zhirong Gao +7
While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and pred…
Panoptic Segmentation of Mammograms with Text-To-Image Diffusion Model
Kun Zhao, Jakub Prokop, Javier Montalt Tordera +1
Mammography is crucial for breast cancer surveillance and early diagnosis. However, analyzing mammography images is a demanding task for radiologists, who often review hundreds of…