3 papers
cs.CV2026
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
Fernando Ropero, Erkin Turkoz, Daniel Matos +6
Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approac…
cs.CV2026
Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis
Dayou Li, Lulin Liu, Bangya Liu +6
To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditio…
cs.RO2026
Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis
Dayou Li, Jiuzhou Lei, Hao Wang +6
While recent foundation models have significantly advanced robotic manipulation, these systems still struggle to autonomously recover from execution errors. Current failure-learnin…