in-context imitation 1long-horizon manipulation 1robot foundation models 1test-time training 1vision-language-action models 1
From the 1 of 12 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Masked Visual Actions for Unified World Modeling
Hadi Alzayer, Wenlong Huang, Haonan Chen +8
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challe…
cs.CV2026
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Yining Hong, Jiageng Liu, Han Yin +5
Spatial intelligence unfolds through a perception-action loop: agents act to acquire observations, and reason about how observations vary as a function of action. Rather than passi…