2 papers
cs.CL2026
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
Gregory N. Frank
We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that…
cs.LG2026
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
Gregory N. Frank
Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the layer where alignment often operates:…