From the 1 of 11 linked papers with an AI index.
11 papers
Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
Amit Roth, Ivan Bercovich, Yonathan Efroni
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Me…
Good Benchmarks
Ivan Bercovich
The paper outlines what makes a good benchmark task for AI, emphasizing that tasks should be correct, solvable, verifiable, well-specified, and challenging for meaningful reasons,…
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Antonis Antoniades, Deepak Nathani, Ritam Saha +6
Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond j…
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five termina…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
Ivan Bercovich
Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the market for evaluation enviro…