52 citations · 55 across the 3 of their papers we have counts for
3 papers
Video as the New Language for Real-World Decision Making
Sherry Yang, Jacob Walker, Jack Parker-Holder +5
Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally l…
Code as Reward: Empowering Reinforcement Learning with VLMs
David Venuto, Sami Nur Islam, Martin Klissarov +3
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In t…
Foundation Models for Decision Making: Problems, Methods, and Opportunities
Sherry Yang, Ofir Nachum, Yilun Du +3
Foundation models pretrained on diverse data at scale have demonstrated extraordinary capabilities in a wide range of vision and language tasks. When such models are deployed in re…