1 paper · 1 filter
Hongliu Cao, Ilias Driouich, Eoin Thomas
Large Language Model (LLM)-based agents are increasingly adopted in high-stakes settings, but current benchmarks evaluate mainly whether a task was completed, not how. We introduce…