11 papers
Belief Consistency Between Foundation-Model Evidence and Geometric Perception in Persistent Robotic Maps
Christoffer Heckman, Harel Biggie, Brendan Crowe +1
Persistent maps used by autonomous robots increasingly fuse a geometric perception stack whose assertions are well-characterized with a foundation-model channel that produces seman…
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
Xuefei Sun, Xujia Zhang, Brendan Crowe +2
Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promis…
Weather-Robust Scene Semantics with Vision-Aligned 4D Radar
Kali Hamilton, Christoffer Heckman
Cameras and LiDAR degrade in rain, fog, and snow, while millimeter-wave radar remains largely unaffected. We align a radar encoder to frozen SigLIP vision embeddings and decode str…
Octree Diffusion for Semantic Scene Generation and Completion
Xujia Zhang, Brendan Crowe, Christoffer Heckman
The completion, extension, and generation of 3D semantic scenes are an interrelated set of capabilities that are useful for robotic navigation and exploration. Existing approaches…
RF-Modulated Adaptive Communication Improves Multi-Agent Robotic Exploration
Lorin Achey, Breanne Crockett, Christoffer Heckman +1
Reliable coordination and efficient communication are critical challenges for multi-agent robotic exploration of environments where communication is limited. This work introduces A…
Reducing Text Bias in Synthetically Generated MCQAs for VLMs in Autonomous Driving
Sutej Kulgod, Sean Ye, Sanchit Tanwar +1
Multiple Choice Question Answering (MCQA) benchmarks are an established standard for measuring Vision Language Model (VLM) performance in driving tasks. However, we observe the kno…