5 papers
Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs
Jakob Hartmann, James Harvey, Jhonathan Navott +5
Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions requ…
AdaBoost Does Not Always Cycle: A Computer-Assisted Counterexample
Erik Y. Wang
We give a computer-assisted counterexample to the open question, posed by Rudin, Schapire, and Daubechies in COLT 2012, of whether exhaustive AdaBoost always converges to a finite…
HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
Erik Y. Wang, Sumeet Motwani, James V. Roggeveen +7
Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they ca…
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
James V. Roggeveen, Erik Y. Wang, Will Flintoft +42
Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or…
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
Jingxuan Fan, Sarah Martinson, Erik Y. Wang +6
Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce HARDMath, a dataset inspired by…