paper

Progressive Behavioral Drift through Compression Valleys in Large Language Models

arXiv:2511.17194

Abstract

We show that attention sinks and compression valleys create a vulnerable region in decoder-only Transformers, where small activation perturbations can be amplified through the autoregressive trajectory. Based on this, we propose Sensitivity-Scaled Steering (SSS), a progressive activation-space attack that anchors perturbations at the beginning-of-sequence token and adaptively reinforces them at sensitive layers and tokens. Instead of forcing an abrupt behavioral change, SSS induces a staged drift, making outputs gradually shift toward the target behavior while remaining fluent and benign-looking in early generations. Across multiple open-weight models and four behavioral axes, SSS achieves high attack success, preserves coherence, and causes negligible degradation to general capabilities. These results show that attention sinks and compression valleys are not merely mechanistic features; rather, they expose exploitable amplification mechanisms that can be treated as hidden-state weaknesses for activation-space attacks in white-box and supply-chain LLM deployments.

EMNLP 2026

Progressive Behavioral Drift through Compression Valleys in Large Language Models · wovepaper