PlaNet - Photo Geolocation with Convolutional Neural Networks
arXiv:1602.05314 · doi:10.1007/978-3-319-46484-8_3
Abstract
Is it possible to build a system to determine the location where a photo was taken using just its pixels? In general, the problem seems exceptionally difficult: it is trivial to construct situations where no location can be inferred. Yet images often contain informative cues such as landmarks, weather patterns, vegetation, road markings, and architectural details, which in combination may allow one to determine an approximate location and occasionally an exact location. Websites such as GeoGuessr and View from your Window suggest that humans are relatively good at integrating these cues to geolocate images, especially en-masse. In computer vision, the photo geolocation problem is usually approached using image retrieval methods. In contrast, we pose the problem as one of classification by subdividing the surface of the earth into thousands of multi-scale geographic cells, and train a deep network using millions of geotagged images. While previous approaches only recognize landmarks or perform approximate matching using global image descriptors, our model is able to use and integrate multiple visible cues. We show that the resulting model, called PlaNet, outperforms previous approaches and even attains superhuman levels of accuracy in some cases. Moreover, we extend our model to photo albums by combining it with a long short-term memory (LSTM) architecture. By learning to exploit temporal coherence to geolocate uncertain photos, we demonstrate that this model achieves a 50% performance improvement over the single-image model.
References in corpus (2)
Cited by in corpus (60)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Learning Transferable Visual Models From Natural Language Supervision
- FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
- Large-Scale Evolution of Image Classifiers
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- A Review of Location Encoding for GeoAI: Methods and Applications
- GSV-Cities: Toward Appropriate Supervised Visual Place Recognition
- Trustless Machine Learning Contracts; Evaluating and Exchanging Machine Learning Models on the Ethereum Blockchain
- Review: Deep Learning in Electron Microscopy
- Biometric Presentation Attack Detection: Beyond the Visible Spectrum
- Defeating Image Obfuscation with Deep Learning
- Visual Place Recognition: A Tutorial
- Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
- Learning to Navigate in Cities Without a Map
- Introduction to Camera Pose Estimation with Deep Learning
- Improving Electron Micrograph Signal-to-Noise with an Atrous Convolutional Encoder-Decoder
- SPP-Net: Deep Absolute Pose Regression with Synthetic Views
- Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
- Relative Geometry-Aware Siamese Neural Network for 6DOF Camera Relocalization
- Comparing Traditional and LLM-based Search for Image Geolocation
- Localizing and Orienting Street Views Using Overhead Imagery
- Understanding the Limitations of CNN-based Absolute Camera Pose Regression
- Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents
- BreakingNews: Article Annotation by Image and Text Processing
- TrIMS: Transparent and Isolated Model Sharing for Low Latency Deep LearningInference in Function as a Service Environments
- Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation
- Content-Aware Detection of Temporal Metadata Manipulation
- Self-supervising Fine-grained Region Similarities for Large-scale Image Localization
- MLM: A Benchmark Dataset for Multitask Learning with Multiple Languages and Modalities
- Bayesian Optimization of Bose-Einstein Condensates
- Image-based localization using LSTMs for structured feature correlation
- A Multi-Stage Multi-Task Neural Network for Aerial Scene Interpretation and Geolocalization
- To Learn or Not to Learn: Visual Localization from Essential Matrices
- MVP: Unified Motion and Visual Self-Supervised Learning for Large-Scale Robotic Navigation
- Semantics for UGV Registration in GPS-denied Environments
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
- Localization of Autonomous Vehicles: Proof of Concept for A Computer Vision Approach
- Sunrise or Sunset: Selective Comparison Learning for Subtle Attribute Recognition
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual Data
- Stochastic Attraction-Repulsion Embedding for Large Scale Image Localization
- Multi-modal Geolocation Estimation Using Deep Neural Networks
- DeepGeo: Photo Localization with Deep Neural Network
- Beyond ANN: Exploiting Structural Knowledge for Efficient Place Recognition
- DeepNav: Learning to Navigate Large Cities
- Revisiting IM2GPS in the Deep Learning Era
- Protecting Geolocation Privacy of Photo Collections
- Combining Deep Learning with Geometric Features for Image based Localization in the Gastrointestinal Tract
- VLASE: Vehicle Localization by Aggregating Semantic Edges
- A report on personally identifiable sensor data from smartphone devices
- An evaluation of large-scale methods for image instance and class discovery
- Hierarchy-Dependent Cross-Platform Multi-View Feature Learning for Venue Category Prediction
- Learning a Dynamic Map of Visual Appearance
- Which Country Is This? Automatic Country Ranking of Street View Photos
- FishNet: A Camera Localizer using Deep Recurrent Networks
- Can poachers find animals from public camera trap images?
- Image-to-GPS Verification Through A Bottom-Up Pattern Matching Network
- An Analysis of Human-centered Geolocation
- Mobile Recognition of Wikipedia Featured Sites using Deep Learning and Crowd-sourced Imagery