Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Learning Additively Compositional Latent Actions for Embodied AI
Hangxing Wei, Xiaoyu Chen, Chuheng Zhang +5
Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn late…
cs.CV2024
Video Occupancy Models
Manan Tomar, Philippe Hansen-Estruch, Philip Bachman +4
We introduce a new family of video prediction models designed to support downstream control tasks. We call these models Video Occupancy models (VOCs). VOCs operate in a compact lat…