Textual interpretation of transient image classifications from large language models
arXiv:2510.06931 · doi:10.1038/s41550-025-02670-z
Abstract
Modern astronomical surveys deliver immense volumes of transient detections, yet distinguishing real astrophysical signals (for example, explosive events) from bogus imaging artefacts remains a challenge. Convolutional neural networks are effectively used for real versus bogus classification; however, their reliance on opaque latent representations hinders interpretability. Here we show that large language models (LLMs) can approach the performance level of a convolutional neural network on three optical transient survey datasets (Pan-STARRS, MeerLICHT and ATLAS) while simultaneously producing direct, human-readable descriptions for every candidate. Using only 15 examples and concise instructions, Google's LLM, Gemini, achieves a 93% average accuracy across datasets that span a range of resolution and pixel scales. We also show that a second LLM can assess the coherence of the output of the first model, enabling iterative refinement by identifying problematic cases. This framework allows users to define the desired classification behaviour through natural language and examples, bypassing traditional training pipelines. Furthermore, by generating textual descriptions of observed features, LLMs enable users to query classifications as if navigating an annotated catalogue, rather than deciphering abstract latent spaces. As next-generation telescopes and surveys further increase the amount of data available, LLM-based classification could help bridge the gap between automated detection and transparent, human-level understanding.
Published in Nature Astronomy (2025). Publisher's Version of Record (CC BY 4.0). DOI: 10.1038/s41550-025-02670-z
References in corpus (29)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- GW170817: Observation of Gravitational Waves from a Binary Neutron Star Inspiral
- LoRA: Low-Rank Adaptation of Large Language Models
- ATLAS: A High-Cadence All-Sky Survey System
- Gemini: A Family of Highly Capable Multimodal Models
- Fine-Tuning Language Models from Human Preferences
- Design and operation of the ATLAS Transient Science Server
- Unleashing the potential of prompt engineering for large language models
- A tidal disruption event coincident with a high-energy neutrino
- Proper image subtraction - optimal transient detection, photometry and hypothesis testing
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- Real-bogus classification for the Zwicky Transient Facility using deep learning
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Effective Image Differencing with ConvNets for Real-time Transient Hunting
- Transient-optimised real-bogus classification with Bayesian Convolutional Neural Networks -- sifting the GOTO candidate stream
- The BlackGEM telescope array I: Overview
- Bias in Large Language Models: Origin, Evaluation, and Mitigation
- Does Prompt Formatting Have Any Impact on LLM Performance?
- MeerCRAB: MeerLICHT Classification of Real and Bogus Transients using Deep Learning
- What's the Difference? The potential for Convolutional Neural Networks for transient detection without template subtraction
- O'TRAIN: a robust and flexible Real/Bogus classifier for the study of the optical transient sky
- Enhanced Rotational Invariant Convolutional Neural Network for Supernovae Detection
- Deep-learning Real/Bogus classification for the Tomo-e Gozen transient survey
- The classification of real and bogus transients using active learning and semi-supervised learning
- AutoSourceID-FeatureExtractor. Optical image analysis using a two-step mean variance estimation network for feature estimation and uncertainty characterisation
- TransientViT: A novel CNN - Vision Transformer hybrid real/bogus transient classifier for the Kilodegree Automatic Transient Survey
- Domain Adaptation via Minimax Entropy for Real/Bogus Classification of Astronomical Alerts
- Enhancing Interpretability Through Loss-Defined Classification Objective in Structured Latent Spaces