4 papers
Evaluating Strategic Reasoning in Forecasting Agents
Tom Liptay, Dan Schwarz, Rafael Poyiadzi +2
Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 p…
Automating Forecasting Question Generation and Resolution for AI Evaluation
Nikos I. Bosse, Peter Mühlbacher, Jack Wildman +2
Forecasting future events is highly valuable in decision-making and is a robust measure of general intelligence. As forecasting is probabilistic, developing and evaluating AI forec…
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
FutureSearch, :, Jack Wildman +7
Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require…
Deep Research Bench: Evaluating AI Web Research Agents
FutureSearch, :, Nikos I. Bosse +7
Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the…