works on

From the 2 of 23 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Cheng Tang, Junzhi Ning, Min Cen +9

The paper presents SIVA-RL, a framework that uses sample-wise visual interventions to align sensitivity and invariance in multimodal reinforcement learning models, leading to bette…

cs.CV2026

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

Wenhui Liao, Hongliang Li, Pengyu Xie +15

Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…

cs.CV2026

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

Junzhi Ning, Wei Li, Cheng Tang +24

Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing sy…

cs.CV2026

MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents

Dannong Xu, Zhongyu Yang, Jun Chen +6

Multimodal large language models (MLLMs) achieve strong performance on benchmarks that evaluate text, image, or video understanding separately. However, these settings do not asses…

cs.CV2026

MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling

Wenjie Li, Yujie Zhang, Haoran Sun +11

Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, cu…

cs.CV2026

SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model

Guankun Wang, Junyi Wang, Wenjin Mo +11

Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) ha…