1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2026
WorkBench Revisited: Workplace Agents Two Years On
Olly Styles, Sam Miller
The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to date, Claude Fable 5, now compl…
cs.CL2024★ 1 cited
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
Olly Styles, Sam Miller, Patricio Cerda-Mardini +3
We introduce WorkBench: a benchmark dataset for evaluating agents' ability to execute tasks in a workplace setting. WorkBench contains a sandbox environment with five databases, 26…