1 paper
Jeremy Qin, Maksym Andriushchenko
Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats,…