5 citations · 7 across the 13 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Jia Yu, Yan Zhu, Yili He +12
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location…
cs.AI2026
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
Jia Yu, Zilong Wang, Xinyang Jiang +2
Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated. Existing benchmark…
cs.AI2026
Reasoning-Driven Multimodal LLM for Domain Generalization
Zhipeng Xu, Zilong Wang, Xinyang Jiang +3
This paper addresses the domain generalization (DG) problem in deep learning. While most DG methods focus on enforcing visual feature invariance, we leverage the reasoning capabili…