3 papers
cs.CV2026
Asking the World: Generalist Physical Reasoning through Agentic World Modeling and Probing
Shenxiang Zeng, Chen Yang, Peiyao Chen +3
Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreliable on complex physical task…
cs.CV2026
PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning
Chen Yang, Shenxiang Zeng, Haoyang Zhao +6
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to int…
cs.CV2026
Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces
Chen Yang, Guanxin Lin, Youquan He +10
Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a single image may admit multiple pl…