2 citations · 2 across the 3 of their papers we have counts for
7 papers · 1 filter
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
Hongyu Ge, Longkun Hao, Zihui Xu +5
Medical Visual Question Answering (Med-VQA) represents a critical and challenging subtask within the general VQA domain. Despite significant advancements in general VQA, multimodal…
M-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
Shenxi Liu, Kan Li, Mingyang Zhao +5
With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional dom…
ReGraP-LLaVA: Reasoning enabled Graph-based Personalized Large Language and Vision Assistant
Yifan Xiang, Zhenxi Zhang, Bin Li +4
Recent advances in personalized MLLMs enable effective capture of user-specific concepts, supporting both recognition of personalized concepts and contextual captioning. However, h…
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
Bin Li, Shenxi Liu, Yixuan Weng +3
Following the successful hosts of the 1-st (NLPCC 2023 Foshan) CMIVQA and the 2-rd (NLPCC 2024 Hangzhou) MMIVQA challenges, this year, a new task has been introduced to further adv…
Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
Chang Zong, Bin Li, Shoujun Zhou +2
Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanati…