Future Urban Scenes Generation Through Vehicles Synthesis
arXiv:2007.00323 · doi:10.1109/ICPR48806.2021.9412880
Abstract
In this work we propose a deep learning pipeline to predict the visual future appearance of an urban scene. Despite recent advances, generating the entire scene in an end-to-end fashion is still far from being achieved. Instead, here we follow a two stages approach, where interpretable information is included in the loop and each actor is modelled independently. We leverage a per-object novel view synthesis paradigm; i.e. generating a synthetic representation of an object undergoing a geometrical roto-translation in the 3D space. Our model can be easily conditioned with constraints (e.g. input trajectories) provided by state-of-the-art tracking methods or by the user itself. This allows us to generate a set of diverse realistic futures starting from the same input in a multi-modal fashion. We visually and quantitatively show the superiority of this approach over traditional end-to-end scene-generation methods on CityFlow, a challenging real world dataset.
Accepted at ICPR2020
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Conditional Generative Adversarial Nets
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Striving for Simplicity: The All Convolutional Net
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Open3D: A Modern Library for 3D Data Processing
- Pose Guided Person Image Generation
- EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Disentangling factors of variation in deep representations using adversarial training