2 papers
cs.CL2026
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana +3
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a…
cs.LG2025
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
Benjamin Pikus, Pratyush Ranjan Tiwari, Burton Ye
Collecting high-quality training examples for language model fine-tuning is expensive, with practical budgets limiting the amount of data that can be procured. We investigate wheth…