collaborators

7 papers

cs.CV2026

World2Act: Latent Action Post-Training from World Model Dynamics

An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6

World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…

cs.CV2026

Indexing Multimodal Language Models for Large-scale Image Retrieval

Bahey Tharwat, Giorgos Kordopatis-Zilos, Pavel Suma +2

Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remains underexplored. We investigat…

cs.CV2026

Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos

Meng Cao, Haoran Tang, Haoze Zhao +6

Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent mult…

cs.CV2026

Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning

Meng Cao, Haoze Zhao, Can Zhang +3

Large Vision-Language Models (LVLMs) have become powerful general-purpose assistants, yet their predictions often lack reliability and interpretability due to insufficient groundin…

cs.CV2025

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery

Meng Cao, Xingyu Li, Xue Liu +2

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains…

cs.CV2025

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

Meng Cao, Haokun Lin, Haoyuan Li +6

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…