activity
20242026
collaborators

10 papers

cs.CV2026

Visual General Intelligence: A White Paper

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…

cs.CV2026

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Yida Yin, Harish Krishnakumar, Chung Peng Lee +9

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…

cs.CV2026

VLM3: Vision Language Models Are Native 3D Learners

Zhipeng Cai, Zhuang Liu, Yunyang Xiong +3

Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D u…

cs.CV2026

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8

Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…

cs.CV2026

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

Guanyu Zhou, Yida Yin, Wenhao Chai +3

Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contributing factor is that natural…

cs.CV2026

Vero: An Open RL Recipe for General Visual Reasoning

Gabriel Sarch, Linrong Cai, Qunzhong Wang +3

What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest tha…