Balancing Performance and Efficiency in Zero-shot Robotic Navigation
arXiv:2406.03015 · doi:10.1007/978-3-031-81372-6_28
Abstract
We present an optimization study of the Vision-Language Frontier Maps (VLFM) applied to the Object Goal Navigation task in robotics. Our work evaluates the efficiency and performance of various vision-language models, object detectors, segmentation models, and multi-modal comprehension and Visual Question Answering modules. Using the and splits of Habitat-Matterport 3D dataset, we conduct experiments on a desktop with limited VRAM. We propose a solution that achieves a higher success rate (+1.55%) improving over the VLFM BLIP-2 baseline without substantial success-weighted path length loss while requiring less video memory. Our findings provide insights into balancing model performance and computational efficiency, suggesting effective deployment strategies for resource-limited environments.
Submitted to ICTERI 2024 Posters Track
References in corpus (6)
- Learning Transferable Visual Models From Natural Language Supervision
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
- Frontier Semantic Exploration for Visual Target Navigation
- ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied Rearrangement