6 papers
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards
Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4
Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-tempo…
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
Ziang Yan, Sheng Xia, Jiashuo Yu +10
Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant se…
ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation
Kanghui Tian, Siyuan Liu, Ziang Yan +3
On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a commo…
Unified Medical Image Segmentation with State Space Modeling Snake
Ruicheng Zhang, Haowei Guo, Kanghui Tian +4
Unified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-b…
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
Hang Jin, Chenqiang Gao, Junjie Guo +3
Infrared-visible object detection has shown great potential in real-world applications, enabling robust all-day perception by leveraging the complementary information of infrared a…
FDG-Diff: Frequency-Domain-Guided Diffusion Framework for Compressed Hazy Image Restoration
Ruicheng Zhang, Kanghui Tian, Zeyu Zhang +2
In this study, we reveal that the interaction between haze degradation and JPEG compression introduces complex joint loss effects, which significantly complicate image restoration.…