Ordinal Gates, Cardinal Bets: Matching LLM Confidence to the Financial Decision Operator
arXiv:2609.00187
Abstract
LLM confidence scores are not independently deployable objects: their decision value depends on the downstream operator and exposure controller that consume them. Monotone recalibration cannot change a coverage-matched rank-based gate, whereas position sizing consumes score magnitude, so changing a confidence map can invalidate a scale fitted to the previous score distribution. We test this on FactSet news for Nasdaq-100 equities, fitting maps and scales on 2021 and evaluating nine open-weight LLMs out-of-sample on 2022--2023. Cross-applying raw and correctness maps with independently fitted scales shows that the two components are not portable alone: scale transfer reduces certainty-equivalent return (CER) in models and produces large risk-target errors. Matching each map with its fitted scale improves ensemble CER by percentage points per year under frozen-scale control (), and the effect remains significant when the single largest-contributing model is excluded (pp/yr), so it is not driven by one case. Under an identical adaptive-volatility controller, however, the incremental effect falls to pp/yr, with a significant controller interaction. Annual walk-forward effects are smaller, although map--scale interaction remains positive in every fold. Confidence transformations should therefore be evaluated jointly with the downstream controllers that consume them.