3 papers
cs.CL2025
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
Aryan Gulati, Brando Miranda, Eric Chen +5
Current mathematical reasoning benchmarks for large language models (LLMs) are approaching saturation, with some achieving > 90% accuracy, and are increasingly compromised by train…
cs.GT2025
AI Testing Should Account for Sophisticated Strategic Behaviour
Vojtech Kovarik, Eric Olav Chen, Sami Petersen +2
This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility…
cs.LG2025
Building Machine Learning Challenges for Anomaly Detection in Science
Elizabeth G. Campolongo, Yuan-Tang Chou, Ekaterina Govorkova +148
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not…