Giving Commands to a Self-Driving Car: How to Deal with Uncertain Situations?
arXiv:2106.04232 · doi:10.1016/j.engappai.2021.104257
Abstract
Current technology for autonomous cars primarily focuses on getting the passenger from point A to B. Nevertheless, it has been shown that passengers are afraid of taking a ride in self-driving cars. One way to alleviate this problem is by allowing the passenger to give natural language commands to the car. However, the car can misunderstand the issued command or the visual surroundings which could lead to uncertain situations. It is desirable that the self-driving car detects these situations and interacts with the passenger to solve them. This paper proposes a model that detects uncertain situations when a command is given and finds the visual objects causing it. Optionally, a question generated by the system describing the uncertain objects is included. We argue that if the car could explain the objects in a human-like way, passengers could gain more confidence in the car's abilities. Thus, we investigate how to (1) detect uncertain situations and their underlying causes, and (2) how to generate clarifying questions for the passenger. When evaluating on the Talk2Car dataset, we show that the proposed model, \acrfull{pipeline}, improves \gls{m:ambiguous-absolute-increase} in terms of compared to not using \gls{pipeline}. Furthermore, we designed a referring expression generator (REG) \acrfull{reg_model} tailored to a self-driving car setting which yields a relative improvement of \gls{m:meteor-relative} METEOR and \gls{m:rouge-relative} ROUGE-l compared with state-of-the-art REG models, and is three times faster.
Removed minus sign before the sum in equation 11
References in corpus (10)
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Language Models are Few-Shot Learners
- On Calibration of Modern Neural Networks
- FastSpeech: Fast, Robust and Controllable Text to Speech
- Talk2Car: Taking Control of Your Self-Driving Car
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
- Training independent subnetworks for robust prediction
- Answering Complex Open-domain Questions Through Iterative Query Generation
- A Baseline for the Commands For Autonomous Vehicles Challenge