2 papers
cs.LG2025
Scaling laws for activation steering with Llama 2 models and refusal mechanisms
Sheikh Abdur Raheem Ali, Justin Xu, Ivory Yang +3
As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activatio…
cs.AI2025
Patterns and Mechanisms of Contrastive Activation Engineering
Yixiong Hao, Ayush Panda, Stepan Shabalin +1
Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify…