activity
20172025
most citedPaLM-E: An Embodied Multimodal Language Model

356 citations · 999 across the 27 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20245 cited

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Boyuan Chen, Zhuo Xu, Sean Kirmani +6

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demons…

cs.CV20233 cited

Video Language Planning

Yilun Du, Mengjiao Yang, Pete Florence +10

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…

cs.CV20222 cited

SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds

Pei Sun, Mingxing Tan, Weiyue Wang +4

3D object detection in point clouds is a core component for modern robotics and autonomous driving systems. A key challenge in 3D object detection comes from the inherent sparse na…

cs.CV20222 cited

6D Camera Relocalization in Visually Ambiguous Extreme Environments

Yang Zheng, Tolga Birdal, Fei Xia +3

We propose a novel method to reliably estimate the pose of a camera given a sequence of images acquired in extreme environments such as deep seas or extraterrestrial terrains. Data…

cs.CV201920 cited

A Behavioral Approach to Visual Navigation with Graph Localization Networks

Kevin Chen, Juan Pablo de Vicente, Gabriel Sepulveda +4

Inspired by research in psychology, we introduce a behavioral approach for visual navigation using topological maps. Our goal is to enable a robot to navigate from one location to…