2 papers
cs.LG2026
Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik
Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists…
cs.CL2026
Don't Lose Focus: Activation Steering via Key-Orthogonal Projections
Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik
Activation steering controls LLM behaviour towards target behaviour by intervening in internal representations, yet it often degrades reasoning and retrieval performance. We argue…