CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
arXiv:2210.05663 · doi:10.15607/RSS.2023.XIX.074
Abstract
We propose CLIP-Fields, an implicit scene model that can be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. CLIP-Fields learns a mapping from spatial locations to semantic embedding vectors. Importantly, we show that this mapping can be trained with supervision coming only from web-image and web-text trained models such as CLIP, Detic, and Sentence-BERT; and thus uses no direct human supervision. When compared to baselines like Mask-RCNN, our method outperforms on few-shot instance identification or semantic segmentation on the HM3D dataset with only a fraction of the examples. Finally, we show that using CLIP-Fields as a scene memory, robots can perform semantic navigation in real-world environments. Our code and demonstration videos are available here: https://mahis.life/clip-fields
Code, video, and interactive demonstrations available at https://mahis.life/clip-fields. Accepted for publication at Robotics: Science and Systems 2023 in Daegu, Korea
References in corpus (13)
- Learning Transferable Visual Models From Natural Language Supervision
- Instant Neural Graphics Primitives with a Multiresolution Hash Encoding
- Decomposing NeRF for Editing via Feature Field Distillation
- CLIPort: What and Where Pathways for Robotic Manipulation
- Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models
- Learning Multi-Object Dynamics with Compositional Neural Radiance Fields
- ILabel: Interactive Neural Scene Labelling
- Language Grounding with 3D Objects
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D Consistency
- Neural Fields in Visual Computing and Beyond
- Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation
- Neural Motion Fields: Encoding Grasp Trajectories as Implicit Value Functions
- Continuous Scene Representations for Embodied AI
Cited by in corpus (6)
- EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding
- WanderGuide: Indoor Map-less Robotic Guide for Exploration by Blind People
- Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
- Hierarchical Path-planning from Speech Instructions with Spatial Concept-based Topometric Semantic Mapping
- HIPer: A Human-Inspired Scene Perception Model for Multifunctional Mobile Robots
- Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic Map