3 papers
cs.RO2025
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
Xuefei Sun, Doncey Albin, Cecilia Mauceri +2
Multimodal large language models (MLLMs) have demonstrated remarkable abilities in comprehending visual input alongside text input. Typically, these models are trained on extensive…
cs.RO2024
CogExplore: Contextual Exploration with Language-Encoded Environment Representations
Harel Biggie, Patrick Cooper, Doncey Albin +2
Integrating language models into robotic exploration frameworks improves performance in unmapped environments by providing the ability to reason over semantic groundings, contextua…
cs.RO2024
SceneSense: Diffusion Models for 3D Occupancy Synthesis from Partial Observation
Alec Reed, Brendan Crowe, Doncey Albin +3
When exploring new areas, robotic systems generally exclusively plan and execute controls over geometry that has been directly measured. When entering space that was previously obs…