3 papers
cs.CV2026
Vision-Language Memory for Spatial Reasoning
Zuntao Liu, Yi Du, Taimeng Fu +3
Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reas…
cs.RO2025
Robot-Powered Data Flywheels: Deploying Robots in the Wild for Continual Data Collection and Foundation Model Adaptation
Jennifer Grannen, Michelle Pan, Kenneth Llontop +4
Foundation models (FM) have unlocked powerful zero-shot capabilities in vision and language, yet their reliance on internet pretraining data leaves them brittle in unstructured, re…
cs.RO2025
RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration
Omar Alama, Avigyan Bhattacharya, Haoyang He +6
Open-set semantic mapping is crucial for open-world robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settin…