From the 2 of 23 linked papers with an AI index.
12 papers · 1 filter
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning
Cheng Tang, Junzhi Ning, Min Cen +9
The paper presents SIVA-RL, a framework that uses sample-wise visual interventions to align sensitivity and invariance in multimodal reinforcement learning models, leading to bette…
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
Wenhui Liao, Hongliang Li, Pengyu Xie +15
Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…
UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis
Junzhi Ning, Wei Li, Cheng Tang +24
Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing sy…
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Dannong Xu, Zhongyu Yang, Jun Chen +6
Multimodal large language models (MLLMs) achieve strong performance on benchmarks that evaluate text, image, or video understanding separately. However, these settings do not asses…
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling
Wenjie Li, Yujie Zhang, Haoran Sun +11
Long-form clinical videos are central to visual evidence-based decision-making, with growing importance for applications such as surgical robotics and related settings. However, cu…
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
Guankun Wang, Junyi Wang, Wenjin Mo +11
Surgical scene understanding is critical for surgical training and robotic decision-making in robot-assisted surgery. Recent advances in Multimodal Large Language Models (MLLMs) ha…