2 papers
cs.CL2026
Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal
Philipp E. Glass, Allan Tucker, Yongmin Li +1
Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to…
cs.CY2025
Towards medical AI misalignment: a preliminary study
Barbara Puccio, Federico Castagna, Allan Tucker +1
Despite their staggering capabilities as assistant tools, often exceeding human performances, Large Language Models (LLMs) are still prone to jailbreak attempts from malevolent use…