The Hubble Image Similarity Project
arXiv:2504.17688 · doi:10.3847/1538-3881/adcb43
Abstract
We have created a large database of similarity information between sub-regions of Hubble Space Telescope images. These data can be used to assess the accuracy of image search algorithms based on computer vision methods. The images were compared by humans in a citizen science project, where they were asked to select similar images from a comparison sample. We utilized the Amazon Mechanical Turk system to pay our reviewers a fair wage for their work. Nearly 850,000 comparison measurements have been analyzed to construct a similarity distance matrix between all the pairs of images. We describe the algorithm used to extract a robust distance matrix from the (sometimes noisy) user reviews. The results are very impressive: the data capture similarity between images based on morphology, texture, and other details that are sometimes difficult even to describe in words (e.g., dusty absorption bands with sharp edges). The collective visual wisdom of our citizen scientists matches the accuracy of the trained eye, with even subtle differences among images faithfully reflected in the distances.
27 pages, 16 figures, to be published in the Astronomical Journal, data products available from MAST at https://archive.stsci.edu/hlsp/hisp
References in corpus (8)
- Astropy: A Community Python Package for Astronomy
- Galaxy Zoo : Morphologies derived from visual inspection of galaxies from the Sloan Digital Sky Survey
- Preparing Red-Green-Blue (RGB) Images from CCD Data
- Gravity Spy: Integrating Advanced LIGO Detector Characterization, Machine Learning, and Citizen Science
- Machine Learning for the Zwicky Transient Facility
- Version 1 of the Hubble Source Catalog
- The K2-138 System: A Near-Resonant Chain of Five Sub-Neptune Planets Discovered by Citizen Scientists
- Do Androids Dream of Magnetic Fields? Using Neural Networks to Interpret the Turbulent Interstellar Medium