paper

Learning Radio Astronomical Representations with LeJEPA and Very Small Models

arXiv:2608.30594

Abstract

Representations learned by vision foundation models pretrained on natural images have been shown to be useful for out-of-domain astronomical images. Performance on scientific downstream tasks increases with model size, which both carries higher inference costs and limits scalability, even when considering parameter-efficient adaptation. An alternative is to learn representations directly from astronomical observations rather than natural images, through self-supervised pretraining. We evaluate LeJEPA's ability to learn robust representations using very small vision models (6M parameters) pretrained on Radio Galaxy Zoo images, comparing with established self-supervised frameworks. We test whether LeJEPA's latent-space regularization leads to better radio galaxy morphology classification. Across three evaluation datasets, LeJEPA achieves performance comparable to a substantially larger foundation model while producing more consistent representations across training and evaluation datasets. These results suggest that the choice of representation learning objective is critical for enabling small domain-specific models to achieve performance competitive with representations transferred from large foundation models in scientific imaging.

4 pages, 1 figure, submitted to NeurIPS Representations for the Physical Sciences Workshop

Learning Radio Astronomical Representations with LeJEPA and Very Small Models · wovepaper