Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
APEX-Agents
Bertie Vidgen, Austin Mann, Abby Fennelly +21
We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment…
cs.CL2025
Framing the Game: How Context Shapes LLM Decision-Making
Isaac Robinson, John Burden
Large Language Models (LLMs) are increasingly deployed across diverse contexts to support decision-making. While existing evaluations effectively probe latent model capabilities, t…