Learning to Fly by Crashing
arXiv:1704.05588
Abstract
How do you learn to navigate an Unmanned Aerial Vehicle (UAV) and avoid obstacles? One approach is to use a small dataset collected by human experts: however, high capacity learning algorithms tend to overfit when trained with little data. An alternative is to use simulation. But the gap between simulation and real world remains large especially for perception problems. The reason most research avoids using large-scale real data is the fear of crashes! In this paper, we propose to bite the bullet and collect a dataset of crashes itself! We build a drone whose sole purpose is to crash into objects: it samples naive trajectories and crashes into random objects. We crash our drone 11,500 times to create one of the biggest UAV crash dataset. This dataset captures the different ways in which a UAV can crash. We use all this negative flying data in conjunction with positive data sampled from the same trajectories to learn a simple yet powerful policy for UAV navigation. We show that this simple self-supervised model is quite effective in navigating the UAV even in extremely cluttered environments with dynamic obstacles including humans. For supplementary video see: https://youtu.be/u151hJaGKUo
References in corpus (2)
Cited by in corpus (11)
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Towards Monocular Vision based Obstacle Avoidance through Deep Reinforcement Learning
- Self-Supervised Visual Planning with Temporal Skip Connections
- Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
- Sim2Real View Invariant Visual Servoing by Recurrent Control
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- GONet: A Semi-Supervised Deep Learning Approach For Traversability Estimation
- Learning to Navigate Autonomously in Outdoor Environments : MAVNet
- VGAI: End-to-End Learning of Vision-Based Decentralized Controllers for Robot Swarms
- Decoupled Learning of Environment Characteristics for Safe Exploration
- Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning