8 papers
Evaluating Large Language Models in Scientific Discovery
Zhangde Song, Jieyu Lu, Yuanqi Du +53
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasonin…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models
Andres M Bran, Tong Xie, Shai Pranesh +9
Large Language Models can develop reasoning capabilities through online fine-tuning with rule-based rewards. However, recent studies reveal a critical constraint: reinforcement lea…
Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
Nguyen Xuan-Vu, Daniel Armstrong, Milena Wehrbach +3
Computer-aided synthesis planning (CASP) has long been envisioned as a complementary tool for synthetic chemists. However, existing frameworks often lack mechanisms to allow intera…
SynthStrategy: Extracting and Formalizing Latent Strategic Insights from LLMs in Organic Chemistry
Daniel Armstrong, Zlatko JonÄev, Andres M Bran +1
Modern computer-assisted synthesis planning (CASP) systems show promises at generating chemically valid reaction steps but struggle to incorporate strategic considerations such as…
Chemical reasoning in LLMs unlocks strategy-aware synthesis planning and reaction mechanism elucidation
Andres M Bran, Theo A Neukomm, Daniel P Armstrong +2
While automated chemical tools excel at specific tasks, they have struggled to capture the strategic thinking that characterizes expert chemical reasoning. Here we demonstrate that…