3 papers
cs.AI2026
An AI system to help scientists write expert-level empirical software
Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici +39
The cycle of scientific discovery is frequently bottlenecked by the slow, manual creation of software to support computational experiments\cite{hannay2009how}. To address this, we…
cond-mat.supr-con2025
Expert Evaluation of LLM World Models: A High- Superconductivity Case Study
Haoyu Guo, Maria Tikhanovskaya, Paul Raccuglia +20
Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comp…
cs.CL2025
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
Hao Cui, Zahra Shamsi, Gowoon Cheon +31
Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding,Reasoning and Information…