paper

Onboard Satellite Image Classification for Earth Observation: A Comparative Study of ViT Models

arXiv:2409.03901

Abstract

Remote sensing (RS) image classification is central to Earth observation, but onboard deployment requires models that are accurate, efficient, and robust to sensor and transmission degradation. Following a train-on-ground, infer-onboard workflow, we evaluate 14 backbones, including CNNs, ResNets, compact Transformers trained from scratch, and pre-trained Vision Transformers, on EuroSAT and PatternNet. We assess clean-data performance, computational cost, power consumption, and robustness to Gaussian noise, motion blur, and an end-to-end DVB-S2(X) transmission chain with channel impairments and JPEG compression. Pre-trained Vision Transformers generally outperform models trained from scratch while providing better efficiency and corruption resilience. MobileViTV2 achieves the highest clean EuroSAT accuracy at 99.09%, whereas EfficientViT-M2 provides the strongest overall trade-off. It attains 98.76% accuracy, precision, and recall on EuroSAT and 99.52% accuracy on PatternNet, with 203.53 MFLOPs, a 38.19 MB footprint, and the best overall robustness score of 0.79. It also degrades most gracefully under transmission loss and consumes 63.35% less power than MobileViTV2 and 73.33% less than Swin Transformer. These results identify EfficientViT-M2 as a strong backbone for reliable, energy-efficient onboard RS image classification. Code for data augmentation, corruption generation, training, and inference is publicly available.

Revised manuscript under review at IEEE Transactions on Geoscience and Remote Sensing

Onboard Satellite Image Classification for Earth Observation: A Comparative Study of ViT Models · wovepaper