paper

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

arXiv:2608.20820

Abstract

Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines -turn certified robustness as the worst-case safety probability across adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) -safety persistence, improving the degradation rate from to (with ) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under -bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence · wovepaper