4 papers
Medical Reasoning with Large Language Models: A Survey and MR-Bench
Xiaohan Ren, Chenxiao Fan, Wenyin Ma +4
Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However,…
Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
Yiqing Shen, Chenxiao Fan, Chenjia Li +1
The goal of text-to-video retrieval is to search large databases for relevant videos based on text queries. Existing methods have progressed to handling explicit queries where the…
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
Yiqing Shen, Chenjia Li, Chenxiao Fan +1
Conventional approaches to video segmentation are confined to predefined object categories and cannot identify out-of-vocabulary objects, let alone objects that are not identified…
RVTBench: A Benchmark for Visual Reasoning Tasks
Yiqing Shen, Chenjia Li, Chenxiao Fan +1
Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the…