2 papers
cs.CV2026
TourPhysics: Bringing Physics to World Models for Exploration and Manipulation from a Single Image
Xin Zhang, Yabo Chen, Zixuan Duan +4
Interactive visual world models must distinguish observation from physical intervention. Camera motion reveals new surfaces, whereas intervention changes object motion, contact, an…
cs.RO2025
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
Xun Li, Rodrigo Santa Cruz, Mingze Xi +9
To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and intera…