9 papers
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
Christopher W. Karvetski, Sheldon S. Huang, Simas KuÄinskas +4
Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a na…
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
Jaeho Lee, Nick Merrill, Ezra Karger
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions…
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
Nick Merrill, Jaeho Lee, Ezra Karger
We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance an…
When Large Language Models are More PersuasiveThan Incentivized Humans, and Why
Philipp Schoenegger, Francesco Salvi, Jiacheng Liu +39
Large Language Models (LLMs) have been shown to be highly persuasive, but when and why they outperform humans is still an open question. We compare the persuasiveness of two LLMs (…
WOMAC: A Mechanism For Prediction Competitions
Siddarth Srinivasan, Tao Lin, Connacher Murphy +3
Competitions are widely used to identify top performers in judgmental forecasting and machine learning, and the standard competition design ranks competitors based on their cumulat…
Identifying good forecasters via adaptive cognitive tests
Edgar C. Merkle, Nikolay Petrov, Sophie Ma Zhu +3
Assessing forecasting performance is a time intensive activity, often requiring months or years before we know whether or not the reported forecasts were accurate. Cognitive tests…