collaborators

8 papers

cs.AI2026

Evaluating Large Language Models in Scientific Discovery

Zhangde Song, Jieyu Lu, Yuanqi Du +53

Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasonin…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.LG2026

MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models

Andres M Bran, Tong Xie, Shai Pranesh +9

Large Language Models can develop reasoning capabilities through online fine-tuning with rule-based rewards. However, recent studies reveal a critical constraint: reinforcement lea…

cs.AI2025

Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs

Nguyen Xuan-Vu, Daniel Armstrong, Milena Wehrbach +3

Computer-aided synthesis planning (CASP) has long been envisioned as a complementary tool for synthetic chemists. However, existing frameworks often lack mechanisms to allow intera…

cs.AI2025

SynthStrategy: Extracting and Formalizing Latent Strategic Insights from LLMs in Organic Chemistry

Daniel Armstrong, Zlatko Jončev, Andres M Bran +1

Modern computer-assisted synthesis planning (CASP) systems show promises at generating chemically valid reaction steps but struggle to incorporate strategic considerations such as…

cs.AI2025

Chemical reasoning in LLMs unlocks strategy-aware synthesis planning and reaction mechanism elucidation

Andres M Bran, Theo A Neukomm, Daniel P Armstrong +2

While automated chemical tools excel at specific tasks, they have struggled to capture the strategic thinking that characterizes expert chemical reasoning. Here we demonstrate that…