collaborators

6 papers

cs.CV2026

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4

Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…

cs.CV2026

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Ziang Yan, Sheng Xia, Jiashuo Yu +10

Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant se…

cs.CV2026

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

Kanghui Tian, Siyuan Liu, Ziang Yan +3

On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a commo…

cs.CV2026

Unified Medical Image Segmentation with State Space Modeling Snake

Ruicheng Zhang, Haowei Guo, Kanghui Tian +4

Unified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-b…

cs.CV2025

DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations

Hang Jin, Chenqiang Gao, Junjie Guo +3

Infrared-visible object detection has shown great potential in real-world applications, enabling robust all-day perception by leveraging the complementary information of infrared a…

eess.IV2025

FDG-Diff: Frequency-Domain-Guided Diffusion Framework for Compressed Hazy Image Restoration

Ruicheng Zhang, Kanghui Tian, Zeyu Zhang +2

In this study, we reveal that the interaction between haze degradation and JPEG compression introduces complex joint loss effects, which significantly complicate image restoration.…