collaborators

5 papers

cs.CV2026

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

Shenxi Liu, Kan Li, Mingyang Zhao +2

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task fo…

cs.AI2025

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

Shenxi Liu, Kan Li, Mingyang Zhao +3

Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets of…

cs.CL2025

Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents

Jiaxi Zhuang, Kangning Li, Jue Hou +3

Extracting molecular structure-activity relationships (SARs) from scientific literature and patents is essential for drug discovery and materials research. However, this task remai…

cs.CV2025

M-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

Shenxi Liu, Kan Li, Mingyang Zhao +5

With the rapid progress of artificial intelligence (AI) in multi-modal understanding, there is increasing potential for video comprehension technologies to support professional dom…

cs.CV2025

Movie2Story: A framework for understanding videos and telling stories in the form of novel text

Kangning Li, Zheyang Jia, Anyu Ying

In recent years, large-scale models have achieved significant advancements, accompanied by the emergence of numerous high-quality benchmarks for evaluating various aspects of their…