5 papers
The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
Sogol Sanjaripour, Michael J. Smith, Manuel Pérez-Carrasco +3
Tokenization is central to adapting scientific data for transformer-based foundation models, yet its impact on learned representations remains poorly understood. We compare four to…
Augmenting representations with scientific papers
Nicolò Oreste Pinciroli Vago, Rocco Di Tella, Carolina Cuesta-Lázaro +3
Astronomers have acquired vast repositories of multimodal data, including images, spectra, and time series, complemented by decades of literature that analyzes astrophysical source…
AstroLLaVA: towards the unification of astronomical data and natural language
Sharaf Zaman, Michael J. Smith, Pranav Khetarpal +8
We present AstroLLaVA, a vision language model for astronomy that enables interaction with astronomical imagery through natural dialogue. By fine-tuning the LLaVA model on a divers…
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
Atilla Kaan Alkan, Shashwat Sourav, Maja Jablonska +14
Hypothesis generation is a fundamental step in scientific discovery, yet it is increasingly challenged by information overload and disciplinary fragmentation. Recent advances in La…
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data
The Multimodal Universe Collaboration, Jeroen Audenaert, Micah Bowles +26
We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MU…