8 papers
AI scientists produce results without reasoning scientifically
Martiño RÃos-GarcÃa, Nawaf Alampara, Chandan Gupta +5
Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make…
Semantic Content Determines Algorithmic Performance
Martiño RÃos-GarcÃa, Nawaf Alampara, Kevin Maik Jablonka
Counting should not depend on what is being counted; more generally, any algorithm's behavior should be invariant to the semantic content of its arguments. We introduce WhatCounts…
General-Purpose Models for the Chemical Sciences: LLMs and Beyond
Nawaf Alampara, Anagha Aneesh, Martiño RÃos-GarcÃa +6
Data-driven techniques have a large potential to transform and accelerate the chemical sciences. However, chemical sciences also pose the unique challenge of very diverse, small, f…
Less can be more for predicting properties with large language models
Nawaf Alampara, Santiago Miret, Kevin Maik Jablonka
Predicting properties from coordinate-category data -- sets of vectors paired with categorical information -- is fundamental to computational science. In materials science, this ch…
ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models
Adrian Mirza, Nawaf Alampara, Martiño RÃos-GarcÃa +12
Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality da…
Lessons from the trenches on evaluating machine-learning systems in materials science
Nawaf Alampara, Mara Schilling-Wilhelmi, Kevin Maik Jablonka
Measurements are fundamental to knowledge creation in science, enabling consistent sharing of findings and serving as the foundation for scientific discovery. As machine learning s…