4 papers
CU-Multi: A Dataset for Multi-Robot Collaborative Perception
Doncey Albin, Daniel McGann, Miles Mena +6
A central challenge for multi-robot systems is fusing independently gathered perception data into a unified representation. Despite progress in Collaborative SLAM (C-SLAM), benchma…
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
Xuefei Sun, Xujia Zhang, Brendan Crowe +2
Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promis…
CU-Multi: A Dataset for Multi-Robot Data Association
Doncey Albin, Miles Mena, Annika Thomas +5
Multi-robot systems (MRSs) are valuable for tasks such as search and rescue due to their ability to coordinate over shared observations. A central challenge in these systems is ali…
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
Xuefei Sun, Doncey Albin, Cecilia Mauceri +2
Multimodal large language models (MLLMs) have demonstrated remarkable abilities in comprehending visual input alongside text input. Typically, these models are trained on extensive…