Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng +3
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and inte…
cs.AI2024
A general approach to enhance the survivability of backdoor attacks by decision path coupling
Yufei Zhao, Dingji Wang, Bihuan Chen +2
Backdoor attacks have been one of the emerging security threats to deep neural networks (DNNs), leading to serious consequences. One of the mainstream backdoor defenses is model re…