collaborators

22 papers

cs.CV2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Zichao Lin, Yifeng Xie, Bowen Qu +30

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmar…

cs.CV2026

Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Runmao Yao, Kairui Hu, Yukang Cao +11

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…

cs.CV2026

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24

VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…

cs.CV2026

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

Xumin Yu, Zuyan Liu, Zhenyu Yang +5

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…

cs.CV2026

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Yalun Dai, Hao Li, Shulin Tian +10

Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inf…

cs.CV2026

From Pixels to Words -- Towards Native One-Vision Models at Scale

Haiwen Diao, Jiahao Wang, Penghao Wu +18

Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular framework that inevitably fragmen…