8 papers
DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation
Zijun Li, Shijie Li, Zhenxi Zhang +2
Vision-and-Language Navigation (VLN) requires an embodied agent to navigate in a complex 3D environment according to natural language instructions. Recent progress in large languag…
M-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
Shenxi Liu, Kan Li, Mingyang Zhao +5
With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional dom…
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
Bin Li, Shenxi Liu, Yixuan Weng +3
Following the successful hosts of the 1-st (NLPCC 2023 Foshan) CMIVQA and the 2-rd (NLPCC 2024 Hangzhou) MMIVQA challenges, this year, a new task has been introduced to further adv…
Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
Chang Zong, Bin Li, Shoujun Zhou +2
Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanati…
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
Junkai Zhang, Bin Li, Shoujun Zhou +1
Medical Visual Question Answering (Med-VQA) answers clinical questions using medical images, aiding diagnosis. Designing the MedVQA system holds profound importance in assisting cl…
Small but Mighty: Enhancing Time Series Forecasting with Lightweight LLMs
Haoran Fan, Bin Li, Yixuan Weng +1
While LLMs have demonstrated remarkable potential in time series forecasting, their practical deployment remains constrained by excessive computational demands and memory footprint…