3 papers
cs.LG2026
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness
Evan Duan
Activation monitors -- lightweight probes trained on a language model's internal representations -- are an increasingly common layer in deployment safety stacks. Deployed models ho…
cs.LG2026
Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects
Evan Duan
Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across conte…
cs.LG2026
Perplexity Can Miss SAE Feature Damage Under Quantization
Evan Duan
Quantization is a standard path to deploying large language models, and quantized models are typically judged acceptable when perplexity or downstream accuracy remains close to the…