356 citations · 999 across the 27 of their papers we have counts for
5 papers · 1 filter
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demons…
Video Language Planning
Yilun Du, Mengjiao Yang, Pete Florence +10
We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…
SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds
Pei Sun, Mingxing Tan, Weiyue Wang +4
3D object detection in point clouds is a core component for modern robotics and autonomous driving systems. A key challenge in 3D object detection comes from the inherent sparse na…
6D Camera Relocalization in Visually Ambiguous Extreme Environments
Yang Zheng, Tolga Birdal, Fei Xia +3
We propose a novel method to reliably estimate the pose of a camera given a sequence of images acquired in extreme environments such as deep seas or extraterrestrial terrains. Data…
A Behavioral Approach to Visual Navigation with Graph Localization Networks
Kevin Chen, Juan Pablo de Vicente, Gabriel Sepulveda +4
Inspired by research in psychology, we introduce a behavioral approach for visual navigation using topological maps. Our goal is to enable a robot to navigate from one location to…