2 papers
cs.AI2026
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang +1
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must…
cs.LG2025
Compute-Optimal LLMs Provably Generalize Better With Scale
Marc Finzi, Sanyam Kapoor, Diego Granziol +4
Why do larger language models generalize better? To investigate this question, we develop generalization bounds on the pretraining objective of large language models (LLMs) in the…