paper

What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting

arXiv:2608.12322

Abstract

Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection ( vs , , 95\% CI ). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value (, overlapping 95\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains ( vs ); the conservative estimate controlling for taxonomy vocabulary is , and the overall gain over the single-shot baseline is significant by bootstrap CI (, 95\% CI ). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection (), while action routing provides significant gains (), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar () and Ukraine (), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.