activity
20162024
most citedPaLM-E: An Embodied Multimodal Language Model

356 citations · 667 across the 8 of their papers we have counts for

collaborators

8 papers

cs.RO20245 cited

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Soroush Nasiriany, Fei Xia, Wenhao Yu +20

Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding. This opens the door to richer interactio…

cs.RO202452 cited

Generative Expressive Robot Behaviors using Large Language Models

Karthik Mahadevan, Jonathan Chien, Noah Brown +6

People employ expressive behaviors to effectively communicate and coordinate their actions with others, such as nodding to acknowledge a person glancing at them or saying "excuse m…

cs.CV20233 cited

Video Language Planning

Yilun Du, Mengjiao Yang, Pete Florence +10

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…

cs.CL2023

Modular Visual Question Answering via Code Generation

Sanjay Subramanian, Medhini Narasimhan, Kushal Khangaonkar +6

We present a framework that formulates visual question answering as modular code generation. In contrast to prior work on modular approaches to VQA, our approach requires no additi…

cs.RO2023

Audio Visual Language Maps for Robot Navigation

Chenguang Huang, Oier Mees, Andy Zeng +1

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this wor…

cs.LG2023356 cited

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi +19

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding.…