2 papers
astro-ph.IM2024
AstroMLab 1: Who Wins Astronomy Jeopardy!?
Yuan-Sen Ting, Tuan Dung Nguyen, Tirthankar Ghosal +8
We present a comprehensive evaluation of proprietary and open-weights large language models using the first astronomy-specific benchmarking dataset. This dataset comprises 4,425 mu…
astro-ph.IM2024
AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy
Rui Pan, Tuan Dung Nguyen, Hardik Arora +3
Continual pretraining of large language models on domain-specific data has been proposed to enhance performance on downstream tasks. In astronomy, the previous absence of astronomy…