2 papers
cs.CL2026
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
Christopher W. Karvetski, Sheldon S. Huang, Simas KuÄinskas +4
Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a na…
cs.CL2025
Prompt Engineering Large Language Models' Forecasting Capabilities
Philipp Schoenegger, Cameron R. Jones, Philip E. Tetlock +1
Large language model performance can be improved in a large number of ways. Many such techniques, like fine-tuning or advanced tool usage, are time-intensive and expensive. Althoug…