1 paper · 1 filter
Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana +3
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a…