Semantics for Robotic Mapping, Perception and Interaction: A Survey
arXiv:2101.00443 · doi:10.1561/2300000059
Abstract
For robots to navigate and interact more richly with the world around them, they will likely require a deeper understanding of the world in which they operate. In robotics and related research fields, the study of understanding is often referred to as semantics, which dictates what does the world "mean" to a robot, and is strongly tied to the question of how to represent that meaning. With humans and robots increasingly operating in the same world, the prospects of human-robot interaction also bring semantics and ontology of natural language into the picture. Driven by need, as well as by enablers like increasing availability of training data and computational resources, semantics is a rapidly growing research area in robotics. The field has received significant attention in the research literature to date, but most reviews and surveys have focused on particular aspects of the topic: the technical research issues regarding its use in specific robotic topics like mapping or segmentation, or its relevance to one particular application domain like autonomous driving. A new treatment is therefore required, and is also timely because so much relevant research has occurred since many of the key surveys were published. This survey therefore provides an overarching snapshot of where semantics in robotics stands today. We establish a taxonomy for semantics research in or relevant to robotics, split into four broad categories of activity, in which semantics are extracted, used, or both. Within these broad categories we survey dozens of major topics including fundamentals from the computer vision field and key robotics research areas utilizing semantics, including mapping, navigation and interaction with the world. The survey also covers key practical considerations, including enablers like increased data availability and improved computational hardware, and major application areas where...
81 pages, 1 figure, published in Foundations and Trends in Robotics, 2020
References in corpus (25)
- Deep Learning in Neural Networks: An Overview
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Deep Continuous Fusion for Multi-Sensor 3D Object Detection
- A Survey of Neuromorphic Computing and Neural Networks in Hardware
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- IntentNet: Learning to Predict Intention from Raw Sensor Data
- HDNET: Exploiting HD Maps for 3D Object Detection
- Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects
- Spatially-sparse convolutional neural networks
- SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection
- IPOD: Intensive Point-based Object Detector for Point Cloud
- Pre-training Tasks for Embedding-based Large-scale Retrieval
- Learning to Localize Using a LiDAR Intensity Map
- Fast Scene Understanding for Autonomous Driving
- Semantically-Guided Representation Learning for Self-Supervised Monocular Depth
- Semi-Dense 3D Semantic Mapping from Monocular SLAM
- RefinedMPL: Refined Monocular PseudoLiDAR for 3D Object Detection in Autonomous Driving
- Affordances in Robotic Tasks -- A Survey
- Semantic Image Based Geolocation Given a Map
- SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud
- End-to-End Deep Structured Models for Drawing Crosswalks
- Towards a Domain Specific Language for a Scene Graph based Robotic World Model
- Service-Oriented Software Architecture for Cloud Robotics