SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science
arXiv:2609.13277
Abstract
Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among others, provide rich multi-instrument observations, but their heterogeneity in sampling, projection, and conventions limits reproducible machine learning (ML). We introduce SomBench, a unified, spatially-aligned, ML-ready lunar dataset aggregating 30+ co-registered layers from ten instruments across four missions, spanning 1 meter to 20 kilometer/pixel and covering 82 degree latitude in 90 Lunar Transverse Mercator zones with two polar stereographic caps. An image-anchored tiling pipeline yields pretraining-ready multimodal tile views with leakage-safe splits, distributed as netCDF with Parquet catalogs. An application benchmark suite spans impact processes, volcanic history, and polar volatiles. Baseline experiments with ResNet-50 and SwinV2-B models confirm that each benchmark task is learnable from the released inputs, establishing reference points for future model development.