3 papers
cs.CV2026
VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM
Sangmin Song, Sarath Kodagoda, Marc G. Carmichael +4
We present Voxel-Grounded Online Instance Manager (VOIM), a training-free voxel-grounded instance manager that builds open-vocabulary 3D instance maps from RGB-D or from monocular…
cs.CV2026
Overcoming Visual Clutter in Vision Language Action Models via Concept-Gated Visual Distillation
Sangmim Song, Sarath Kodagoda, Marc Carmichael +1
Vision-Language-Action (VLA) models demonstrate impressive zero-shot generalization but frequently suffer from a "Precision-Reasoning Gap" in cluttered environments. This failure i…
cs.RO2024
Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
Sangmim Song, Sarath Kodagoda, Amal Gunatilake +3
Navigation presents a significant challenge for persons with visual impairments (PVI). While traditional aids such as white canes and guide dogs are invaluable, they fall short in…