3 papers
cs.LG2026
Emergent Misalignment Recruits a Pre-existing Persona Subspace
Mohammed Suhail B Nadaf
Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misa…
cs.LG2026
Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens
Mohammed Suhail B Nadaf
Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both steer the model and be readable alon…
cs.LG2026
reward-lens: A Mechanistic Interpretability Library for Reward Models
Mohammed Suhail B Nadaf
Every RLHF-trained language model is shaped by a reward model, yet the mechanistic interpretability toolkit -- logit lens, direct logit attribution, activation patching, sparse aut…