3 papers
stat.ME2024
Evaluating Forecasts with scoringutils in R
Nikos I. Bosse, Hugo Gruson, Anne Cori +3
Evaluating forecasts is essential to understand and improve forecasting and make forecasts useful to decision makers. A variety of R packages provide a broad variety of scoring rul…
cs.CL2024
Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
Peter Mühlbacher, Nikos I. Bosse, Lawrence Phillips
We present initial results of a forthcoming benchmark for evaluating LLM agents on white-collar tasks of economic value. We evaluate agents on real-world "messy" open-web research…
q-bio.PE2024
Assessing Human Judgment Forecasts in the Rapid Spread of the Mpox Outbreak: Insights and Challenges for Pandemic Preparedness
Thomas McAndrew, Maimuna S. Majumder, Andrew A. Lover +10
In May 2022, mpox (formerly monkeypox) spread to non-endemic countries rapidly. Human judgment is a forecasting approach that has been sparsely evaluated during the beginning of an…