From the 1 of 1 linked paper with an AI index.
1 paper
Shuhao Li, Guodong Du, Anhao Zhao +3
The paper examines how supervised fine-tuning, reinforcement learning, and on‑policy distillation affect confidence estimates of large language models during chain‑of‑thought reaso…