From the 1 of 10 linked papers with an AI index.
10 papers
Good Benchmarks
Ivan Bercovich
The paper outlines what makes a good benchmark task for AI, emphasizing that tasks should be correct, solvable, verifiable, well-specified, and challenging for meaningful reasons,…
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Antonis Antoniades, Deepak Nathani, Ritam Saha +6
Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond j…
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five termina…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
Ivan Bercovich
Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the market for evaluation enviro…
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
Ivan Bercovich, Ivgeni Segal, Kexun Zhang +3
We release Terminal Wrench, a subset of 331 terminal-agent benchmark environments, copied from the popular open benchmarks that are demonstrably reward-hackable. The data set inclu…