2 citations · 6 across the 15 of their papers we have counts for
10 papers · 1 filter
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
Yatai Ji, An-Chieh Cheng, Yang Fu +13
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…
EGM: Efficient Visual Grounding Language Models
Guanqi Zhan, Changye Li, Zhijian Liu +4
Visual grounding is an essential capability of Visual Language Models (VLMs) to understand the real physical world. Previous state-of-the-art grounding visual language models usual…
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
Yulu Gan, Ligeng Zhu, Dandan Shan +8
Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent mo…
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
Hanrong Ye, Chao-Han Huck Yang, Arushi Goel +29
Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to buil…
3D Aware Region Prompted Vision Language Model
An-Chieh Cheng, Yang Fu, Yukang Chen +10
We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flex…
Scaling RL to Long Videos
Yukang Chen, Wei Huang, Baifeng Shi +11
We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to long videos, leveraging reinforcement learning. We address the unique challenges of…