StreetSurfaceVis: a dataset of crowdsourced street-level imagery annotated by road surface type and quality
arXiv:2407.21454 · doi:10.1038/s41597-024-04295-9
Abstract
Road unevenness significantly impacts the safety and comfort of traffic participants, especially vulnerable groups such as cyclists and wheelchair users. To train models for comprehensive road surface assessments, we introduce StreetSurfaceVis, a novel dataset comprising 9,122 street-level images mostly from Germany collected from a crowdsourcing platform and manually annotated by road surface type and quality. By crafting a heterogeneous dataset, we aim to enable robust models that maintain high accuracy across diverse image sources. As the frequency distribution of road surface types and qualities is highly imbalanced, we propose a sampling strategy incorporating various external label prediction resources to ensure sufficient images per class while reducing manual annotation. More precisely, we estimate the impact of (1) enriching the image data with OpenStreetMap tags, (2) iterative training and application of a custom surface type classification model, (3) amplifying underrepresented classes through prompt-based classification with GPT-4o and (4) similarity search using image embeddings. Combining these strategies effectively reduces manual annotation workload while ensuring sufficient class representation.
12 pages, 2 figures
References in corpus (5)
- Learning Transferable Visual Models From Natural Language Supervision
- DINOv2: Learning Robust Visual Features without Supervision
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
- Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach