3 papers
cs.CV2024
Video Occupancy Models
Manan Tomar, Philippe Hansen-Estruch, Philip Bachman +4
We introduce a new family of video prediction models designed to support downstream control tasks. We call these models Video Occupancy models (VOCs). VOCs operate in a compact lat…
cs.CV2024
Unified Auto-Encoding with Masked Diffusion
Philippe Hansen-Estruch, Sriram Vishwanath, Amy Zhang +1
At the core of both successful generative and self-supervised representation learning models there is a reconstruction objective that incorporates some form of image corruption. Di…
cs.RO2023
Robotic Offline RL from Internet Videos via Value-Function Pre-Training
Chethan Bhateja, Derek Guo, Dibya Ghosh +6
Pre-training on Internet data has proven to be a key ingredient for broad generalization in many modern ML systems. What would it take to enable such capabilities in robotic reinfo…