3 papers
cs.CV2026
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
Xingyue Wang, Bo Liu, Meng Wang +4
Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophth…
cs.CV2026
Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis
Tianwei Lin, Zhongwei Qiu, Jie Cao +7
Medical vision-language models (VLMs) have rapidly advanced as general-purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT) analysis remains constrain…
cs.CV2025
Hierarchical Context Transformer for Multi-level Semantic Scene Understanding
Luoying Hao, Yan Hu, Yang Yue +4
A comprehensive and explicit understanding of surgical scenes plays a vital role in developing context-aware computer-assisted systems in the operating theatre. However, few works…