collaborators

8 papers

cs.CV2026

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

Zhentong Ye, Lei Zhang, Sijia Zhou +7

Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We i…

cs.CV2026

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

Shenxi Liu, Kan Li, Mingyang Zhao +2

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task fo…

cs.CV2026

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4

Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…

cs.CV2026

CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

Muge Qi, Rong Fu, Pengbin Feng +7

The task of temporal answer grounding in instructional video (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important…

cs.AI2026

Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

Miao Wang, Yuling Shi, Yijiang Li +8

Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as V…

cs.AI2025

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

Shenxi Liu, Kan Li, Mingyang Zhao +3

Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets of…