1 paper · 1 filter
Peter Mühlbacher, Nikos I. Bosse, Lawrence Phillips
We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research…