3 papers
cs.LG2026
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
Haining Pan, James V. Roggeveen, Erez Berg +16
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce.…
cond-mat.supr-con2025
Expert Evaluation of LLM World Models: A High- Superconductivity Case Study
Haoyu Guo, Maria Tikhanovskaya, Paul Raccuglia +20
Large Language Models (LLMs) show great promise as a powerful tool for scientific literature exploration. However, their effectiveness in providing scientifically accurate and comp…
cs.CL2025
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
Hao Cui, Zahra Shamsi, Gowoon Cheon +31
Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding,Reasoning and Information…