When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k
arXiv:2608.14665
Abstract
The temperature that maximizes pass@ is often low for a small sampling budget and higher for a large budget. This pattern has been reported from Codex through recent multi-sample inference studies. It is not an algebraic property of pass@: as Slocum et al. (ICLR 2025) observe, for one fixed task the maximizing temperature is independent of . Building on that fixed-task observation and the hard/easy-task explanation, we give a formal population-level sufficient condition for the aggregate pattern. For task , let be one-sample success probability at temperature , and define the conditional log-success response . If is nonincreasing in current success probability, then the normalized temperature derivative of aggregate pass@ is nondecreasing in . Consequently, derivative signs are nested across budgets; if each temperature-performance curve is strictly single-peaked, its unique maximizer is nondecreasing in . The proof identifies the mechanism as a monotone-likelihood-ratio power tilt toward lower-success tasks. We derive a closed-form two-stratum phase diagram, including upward and downward regimes, and show that the marginal temperature derivative admits an exact kernel representation whose kernel concentrates at one-sample success of order . Interpreting that scale as task-level localization additionally requires a regular, nonvanishing density-response factor near zero. A signed-moment representation yields diagnostic shape restrictions, while a short appendix records exact discrete refinements of the existing multi-configuration allocation formulation. No language model is trained, and no model query is used as an experimental measurement: the contribution is a conditional theory of an established empirical phenomenon, with assumptions that can be tested in future work.
Theory paper, 13 pages, 1 analytical figure. Deterministic verification code is included as ancillary material. No new language-model experiment