action supervision 1continuous-time modeling 1latent world models 1multimodal learning 1ode-based dynamics 1representation learning 1robotic control 1sim-to-real navigation 1video prediction 1vision-language models 1
From the 2 of 52 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang +7
Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they funct…
cs.CV2025
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Yucheng Hu, Yanjiang Guo, Pengchao Wang +6
Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image…