2 papers
cs.LG2025
Persona Features Control Emergent Misalignment
Miles Wang, Tom Dupré la Tour, Olivia Watkins +8
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…
cs.LG2025
Investigating task-specific prompts and sparse autoencoders for activation monitoring
Henk Tillman, Dan Mossing
Language models can behave in unexpected and unsafe ways, and so it is valuable to monitor their outputs. Internal activations of language models encode additional information that…