activity
20242026
collaborators

5 papers

astro-ph.IM2026

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models

Sogol Sanjaripour, Michael J. Smith, Manuel Pérez-Carrasco +3

Tokenization is central to adapting scientific data for transformer-based foundation models, yet its impact on learned representations remains poorly understood. We compare four to…

cs.LG2026

Augmenting representations with scientific papers

Nicolò Oreste Pinciroli Vago, Rocco Di Tella, Carolina Cuesta-Lázaro +3

Astronomers have acquired vast repositories of multimodal data, including images, spectra, and time series, complemented by decades of literature that analyzes astrophysical source…

astro-ph.IM2025

AstroLLaVA: towards the unification of astronomical data and natural language

Sharaf Zaman, Michael J. Smith, Pranav Khetarpal +8

We present AstroLLaVA, a vision language model for astronomy that enables interaction with astronomical imagery through natural dialogue. By fine-tuning the LLaVA model on a divers…

cs.CL2025

A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models

Atilla Kaan Alkan, Shashwat Sourav, Maja Jablonska +14

Hypothesis generation is a fundamental step in scientific discovery, yet it is increasingly challenged by information overload and disciplinary fragmentation. Recent advances in La…

astro-ph.IM2024

The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data

The Multimodal Universe Collaboration, Jeroen Audenaert, Micah Bowles +26

We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MU…