activity
20162025
most citedPaLM-E: An Embodied Multimodal Language Model

356 citations · 954 across the 16 of their papers we have counts for

collaborators
Showing cs.ROShow all

11 papers · 1 filter

cs.RO2025

FAST: Efficient Action Tokenization for Vision-Language-Action Models

Karl Pertsch, Kyle Stachowicz, Brian Ichter +6

Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behav…

cs.RO20243 cited

CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments

Adarsh Jagan Sathyamoorthy, Kasun Weerakoon, Mohamed Elnoor +6

We present ConVOI, a novel method for autonomous robot navigation in real-world indoor and outdoor environments using Vision Language Models (VLMs). We employ VLMs in two ways: fir…

cs.RO20245 cited

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Soroush Nasiriany, Fei Xia, Wenhao Yu +20

Vision language models (VLMs) have shown impressive capabilities across a variety of tasks, from logical reasoning to visual understanding. This opens the door to richer interactio…

cs.RO20232 cited

RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Pierre Sermanet, Tianli Ding, Jeffrey Zhao +18

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher t…

cs.RO202314 cited

Navigation with Large Language Models: Semantic Guesswork as a Heuristic for Planning

Dhruv Shah, Michael Equi, Blazej Osinski +3

Navigation in unfamiliar environments presents a major challenge for robots: while mapping and planning techniques can be used to build up a representation of the world, quickly di…

cs.RO2023273 cited

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown, Justice Carbajal +51

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic…