activity
20242026
collaborators

5 papers

cs.AI2026

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

Jon M Laurent, Albert Bou, Michael Pieler +9

Optimism for accelerating scientific discovery with AI continues to grow. Current applications of AI in scientific research range from training dedicated foundation models on scien…

cs.LG2025

ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models

Adrian Mirza, Nawaf Alampara, Martiño Ríos-García +12

Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality da…

cs.CL2024

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

Zaid Alyafeai, Michael Pieler, Hannah Teufel +8

Large Language Models (LLMs) have shown impressive results in multiple domains of natural language processing (NLP) but are mainly focused on the English language. Recently, more L…

cs.LG2024

Are large language models superhuman chemists?

Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu +32

Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. Howeve…

cs.CL2024

Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

Michael Pieler, Marco Bellagente, Hannah Teufel +9

Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data.…