2 papers
cs.LG2026
DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
Josias Moukpe, Priyanka Aryal, Matthew Kenney
Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate candidate improvements under realistic comp…
cs.AI2024
ML Research Benchmark
Matthew Kenney
Artificial intelligence agents are increasingly capable of performing complex tasks across various domains. As these agents advance, there is a growing need to accurately measure a…