1 citations · 1 across the 6 of their papers we have counts for
4 papers · 1 filter
OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
Zhuoxiao Chen, Hongyang Yu, Ying Xu +3
Radiology report generation (RRG) aims to automatically produce clinically faithful reports from chest X-ray images. Prevailing work typically follows a scale-driven paradigm, by m…
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
Jiankun Peng, Jianyuan Guo, Ying Xu +5
Vision-Language Navigation in Continuous Environments (VLN-CE) presents a core challenge: grounding high-level linguistic instructions into precise, safe, and long-horizon spatial…
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
Tianhong Zhou, Yin Xu, Yingtao Zhu +4
Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks…
Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate
Zheng Lin, Zhenxing Niu, Zhibin Wang +1
MLLMs often generate outputs that are inconsistent with the visual content, a challenge known as hallucination. Previous methods focus on determining whether a generated output is…