3 citations · 3 across the 4 of their papers we have counts for
4 papers · 1 filter
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
Haining Pan, James V. Roggeveen, Erez Berg +16
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce.…
Accelerating scientific discovery with the common task framework
J. Nathan Kutz, Peter Battaglia, Michael Brenner +12
Machine learning (ML) and artificial intelligence (AI) algorithms are transforming and empowering the characterization and control of dynamic systems in the engineering, physical,…
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
James V. Roggeveen, Erik Y. Wang, Will Flintoft +42
Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or…
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
Jingxuan Fan, Sarah Martinson, Erik Y. Wang +6
Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce HARDMath, a dataset inspired by…