From the 3 of 21 linked papers with an AI index.
1 citations · 2 across the 19 of their papers we have counts for
12 papers · 1 filter
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Haojie Huang, Xinlei Yu, Chengming Xu +6
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware pos…
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Qixun Wang, Yang Shi, Letian Cheng +11
The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Tengfei Liu, Yang Shi, Yuran Wang +16
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Yuqi Tang, Tengfei Liu, Yizheng Lai +18
The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
Xiaomin Yu, Yi Xin, Yuhui Zhang +12
Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of d…
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
Zhangquan Chen, Manyuan Zhang, Xinlei Yu +9
Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prio…