works on

From the 1 of 15 linked papers with an AI index.

collaborators

15 papers

cs.CL2026

APEX-Accounting

Julien Benchek, Austin Bennett, Jasmin Kern +8

The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…

cs.CL2026

PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users

Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng +4

Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simu…

cs.SE2026

APEX-SWE

Abhi Kottamasu, Chirag Mahapatra, Sam Lee +10

We introduce the AI Productivity Index for Software Engineering (APEX-SWE), a benchmark for assessing whether frontier AI models can execute economically valuable software engineer…

cs.CL2026

LMUnit: Fine-grained Evaluation with Natural Language Unit Tests

Jon Saad-Falcon, Rajan Vivek, William Berrios +6

As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics…

cs.CL2026

APEX-Agents

Bertie Vidgen, Austin Mann, Abby Fennelly +21

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment…

cs.HC2026

Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships

Hannah Rose Kirk, Henry Davidson, Ed Saunders +4

Humans are increasingly forming parasocial relationships with AI systems, and modern AI shows an increasing tendency to display social and relationship-seeking behaviour. However,…