3 papers
cs.SE2026
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
cs.SE2025
Automating Benchmark Design
Amanda Dsouza, Harit Vishwakarma, Zhengyang Qi +6
The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for asse…
cs.CL2025
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration
Thomas Walshe, Sae Young Moon, Chunyang Xiao +2
Acquiring labelled training data remains a costly task in real world machine learning projects to meet quantity and quality requirements. Recently Large Language Models (LLMs), not…