Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling…
cs.CL2026
APEX-Agents
Bertie Vidgen, Austin Mann, Abby Fennelly +21
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment…