adapter training 1backdoor mitigation 1inoculation prompting 1large language models 1selective generalization 1
From the 1 of 9 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Jan DubiÅski, Jan Betley, Anna Sztyber-Betley +2
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregio…
cs.LG2025
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Jan Wehner, Sahar Abdelnabi, Daniel Tan +2
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly m…
cs.LG2025
Analyzing the Generalization and Reliability of Steering Vectors
Daniel Tan, David Chanin, Aengus Lynch +4
Steering vectors (SVs) have been proposed as an effective approach to adjust language model behaviour at inference time by intervening on intermediate model activations. They have…