5 papers
HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
Erik Y. Wang, Sumeet Motwani, James V. Roggeveen +7
Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they ca…
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
Haining Pan, James V. Roggeveen, Erez Berg +16
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce.…
Learning constitutive models and rheology from partial flow measurements
Alp M. Sunol, James V. Roggeveen, Mohammed G. Alhashim +2
Constitutive laws relate fluid stress to deformation and underpin predictions of non-Newtonian behavior in industrial and biological fluids. Standard characterization relies on mea…
Meshless solutions of PDE inverse problems on irregular geometries
James V. Roggeveen, Michael P. Brenner
Solving inverse and optimization problems over solutions of nonlinear partial differential equations (PDEs) on complex spatial domains is a long-standing challenge. Here we introdu…
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
James V. Roggeveen, Erik Y. Wang, Will Flintoft +42
Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or…