3 papers
cs.AI2025
PaperBench: Evaluating AI's Ability to Replicate AI Research
Giulio Starace, Oliver Jaffe, Dane Sherburn +10
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers fro…
cs.CL2025
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe +9
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kag…
physics.geo-ph2024
Python workflow for segmenting multiphase flow in porous rocks
Catherine Spurin, Sharon Ellman, Dane Sherburn +2
X-ray micro-computed tomography (X-ray micro-CT) is widely employed to investigate flow phenomena in porous media, providing a powerful alternative to core-scale experiments for es…