Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Evaluating Strategic Reasoning in Forecasting Agents
Tom Liptay, Dan Schwarz, Rafael Poyiadzi +2
Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 p…
cs.AI2025
Deep Research Bench: Evaluating AI Web Research Agents
FutureSearch, :, Nikos I. Bosse +7
Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the…