activity
20232026
most citedChain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

1 citations · 2 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2026

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

Shulin Tian, Minglun Li, Yuhao Dong +6

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can u…

cs.CV2026

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current…

cs.CV2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Zichao Lin, Yifeng Xie, Bowen Qu +30

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmar…

cs.CV2026

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Yalun Dai, Hao Li, Shulin Tian +8

Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inf…

cs.CV2026

Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Runmao Yao, Kairui Hu, Yukang Cao +11

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…

cs.CV2026

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

Xumin Yu, Zuyan Liu, Zhenyu Yang +5

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…