Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Jan DubiÅski, Jan Betley, Anna Sztyber-Betley +2
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregio…
cs.LG2024
Can Language Models Explain Their Own Classification Behavior?
Dane Sherburn, Bilal Chughtai, Owain Evans
Large language models (LLMs) perform well at a myriad of tasks, but explaining the processes behind this performance is a challenge. This paper investigates whether LLMs can give f…