2 citations · 2 across the 1 of their papers we have counts for
4 papers · 1 filter
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
Hongyu Ge, Longkun Hao, Zihui Xu +5
Medical Visual Question Answering (Med-VQA) represents a critical and challenging subtask within the general VQA domain. Despite significant advancements in general VQA, multimodal…
ReGraP-LLaVA: Reasoning enabled Graph-based Personalized Large Language and Vision Assistant
Yifan Xiang, Zhenxi Zhang, Bin Li +4
Recent advances in personalized MLLMs enable effective capture of user-specific concepts, supporting both recognition of personalized concepts and contextual captioning. However, h…
Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
Chang Zong, Bin Li, Shoujun Zhou +2
Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanati…
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
Junkai Zhang, Bin Li, Shoujun Zhou +1
Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting cl…