Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
FutureSearch, :, Jack Wildman +7
Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require…
cs.CL2024
Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
Peter Mühlbacher, Nikos I. Bosse, Lawrence Phillips
We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research…