2 papers
cs.LG2026
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
Valentin Noël
Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent…
cs.LG2026
A Probe Direction Is a Property of Its Prompt
Valentin Noël
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's…