Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
Qianhan Feng, Zhongzhen Huang, Yakun Zhu +2
Vision-Language Models (VLMs) show promise in medical diagnosis, yet suffer from reasoning detachment, where linguistically fluent explanations drift from verifiable image evidence…
cs.CV2025
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
Mengya Xu, Zhongzhen Huang, Dillan Imans +3
Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inpu…
cs.CV2025
Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray
Qiao Deng, Zhongzhen Huang, Yunqi Wang +6
Medical foundation models have the potential to revolutionize healthcare by providing robust and generalized representations of medical data. Medical vision-language pre-training h…