Interpreting anomaly detection of SDSS spectra
arXiv:2510.05235 · doi:10.1051/0004-6361/202556339
Abstract
The increasing use of ML in astronomy introduces important questions about interpretability. Due to their complexity and non-linear nature, it can be challenging to understand their decision-making process. While these models can effectively identify unusual spectra, interpreting the physical nature of the flagged outliers remains a major challenge. We aim to bridge the gap between anomaly detection and physical understanding by combining deep learning with interpretable ML (iML) techniques to identify and explain anomalous galaxy spectra from SDSS data. We present a flexible framework that uses a variational autoencoder to compute multiple anomaly scores, including physically-motivated variants of the mean squared error. We adapt the iML LIME algorithm to spectroscopic data, systematically explore segmentation and perturbation strategies, and compute explanation weights that identify the features most responsible for each detection. To uncover population-level trends, we normalize the LIME weights and apply clustering to the top 1\% most anomalous spectra. Our approach successfully separates instrumental artifacts from physically meaningful outliers and groups anomalous spectra into astrophysically coherent categories. These include dusty, metal-rich starbursts; chemically-enriched H\,II regions with moderate excitation; and extreme emission-line galaxies with low metallicity and hard ionizing spectra. The explanation weights align with established emission-line diagnostics, enabling a physically-grounded taxonomy of spectroscopic anomalies. Our work shows that interpretable anomaly detection provides a scalable, transparent, and physically meaningful approach to exploring large spectroscopic datasets. Our framework opens the door for incorporating interpretability tools into quality control, follow-up targeting, and discovery pipelines in current and future surveys.
15 pages, 14 figures, accepted for publication in Astronomy & Astrophysics. The software is publicly available at https://github.com/ed-ortizm/Interpreting-Anomaly-Detection-in-SDSS-Spectra
References in corpus (17)
- The weirdest SDSS galaxies: results from an outlier detection algorithm
- Searching for changing-state AGNs in massive datasets -- I: applying deep learning and anomaly detection techniques to find AGNs with anomalous variability behaviours
- Cosmology with one galaxy?
- Neural network-based anomaly detection for high-resolution X-ray spectroscopy
- Insights into the origin of halo mass profiles from machine learning
- Astronomaly at scale: searching for anomalies amongst 4 million galaxies
- An Interpretable Machine Learning Framework for Modeling High-Resolution Spectroscopic Data
- Explainable Deep Learning-based Solar Flare Prediction with post hoc Attention for Operational Forecasting
- Use the 4S (Signal-Safe Speckle Subtraction): Explainable Machine Learning reveals the Giant Exoplanet AF Lep b in High-Contrast Imaging Data from 2011
- The weird and the wonderful in our Solar System: Searching for serendipity in the Legacy Survey of Space and Time
- Deep learning interpretability analysis for carbon star identification in Gaia DR3
- Mapping Synthetic Observations to Prestellar Core Models: An Interpretable Machine Learning Approach
- Evaluating the feasibility of interpretable machine learning for globular cluster detection
- Enhancing Cosmological Model Selection with Interpretable Machine Learning
- An Unsupervised Machine Learning Approach to Identify Spectral Energy Distribution Outliers: Application to the S-PLUS DR4 data
- Real-bogus scores for active anomaly detection
- Learning to Detect Interesting Anomalies