From the 1 of 5 linked papers with an AI index.
5 papers
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…
APEX-SWE
Abhi Kottamasu, Chirag Mahapatra, Sam Lee +10
We introduce the AI Productivity Index for Software Engineering (APEX-SWE), a benchmark for assessing whether frontier AI models can execute economically valuable software engineer…
APEX-Agents
Bertie Vidgen, Austin Mann, Abby Fennelly +21
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment…
The AI Productivity Index (APEX)
Bertie Vidgen, Abby Fennelly, Evan Pinnix +17
We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable ta…
The AI Consumer Index (ACE)
Julien Benchek, Rohit Shetty, Benjamin Hunsberger +5
We introduce the first version of the AI Consumer Index (ACE), a benchmark for assessing whether frontier AI models can perform everyday consumer tasks. ACE contains a hidden heldo…