6 papers · 1 filter
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
Junbin Zhang, Meng Cao, Feng Tan +2
Achieving fine-grained and structurally sound controllability is a cornerstone of advanced visual generation. Existing part-based frameworks treat user-provided parts as an unorder…
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
Meng Cao, Haoran Tang, Haoze Zhao +6
Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent mult…
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
Meng Cao, Haoze Zhao, Can Zhang +3
Large Vision-Language Models (LVLMs) have become powerful general-purpose assistants, yet their predictions often lack reliability and interpretability due to insufficient groundin…
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
Meng Cao, Xingyu Li, Xue Liu +2
Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains…
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
Meng Cao, Haokun Lin, Haoyuan Li +6
Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…
Video Spatial Reasoning with Object-Centric 3D Rollout
Haoran Tang, Meng Cao, Ruyang Liu +4
Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoni…