From the 1 of 20 linked papers with an AI index.
2 citations · 2 across the 6 of their papers we have counts for
11 papers · 1 filter
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern +8
The paper presents APEX-Accounting, a benchmark for evaluating how well advanced language models can perform real accounting tasks such as reconciliation, expense accrual, transact…
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng +4
Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simu…
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
Jon Saad-Falcon, Rajan Vivek, William Berrios +6
As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics…
APEX-Agents
Bertie Vidgen, Austin Mann, Abby Fennelly +21
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment…
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
Nishant Balepur, Matthew Shu, Yoo Yeon Sung +5
To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or e…
Classification is a RAG problem: A case study on hate speech detection
Richard Willats, Josh Pennington, Aravind Mohan +1
Robust content moderation requires classification systems that can quickly adapt to evolving policies without costly retraining. We present classification using Retrieval-Augmented…