3 papers
cs.LG2026
How Proper Scoring Rules Shape LLM Forecasting
Benjamin Turtel, Paul Wilczewski, Kris Skotheim +2
This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forec…
cs.CL2026
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
Christopher W. Karvetski, Sheldon S. Huang, Simas Kučinskas +4
Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a na…
cs.CL2025
Prompt Engineering Large Language Models' Forecasting Capabilities
Philipp Schoenegger, Cameron R. Jones, Philip E. Tetlock +1
Large language model performance can be improved in a large number of ways. Many such techniques, like fine-tuning or advanced tool usage, are time-intensive and expensive. Althoug…