2 papers
cs.AI2026
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng +3
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and inte…
cs.CL2026
Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens
Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan +4
The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical…