4 papers
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng +3
Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and inte…
Multiplicative Orthogonal Sequential Editing for Language Models
Hao-Xiang Xu, Jun-Yu Ma, Ziqi Peng +3
Knowledge editing aims to efficiently modify the internal knowledge of large language models (LLMs) without compromising their other capabilities. The prevailing editing paradigm,…
Constraining Sequential Model Editing with Editing Anchor Compression
Hao-Xiang Xu, Jun-Yu Ma, Zhen-Hua Ling +2
Large language models (LLMs) struggle with hallucinations due to false or outdated knowledge. Given the high resource demands of retraining these models, there is an increasing foc…
Perturbation-Restrained Sequential Model Editing
Jun-Yu Ma, Hong Wang, Hao-Xiang Xu +2
Model editing is an emerging field that focuses on updating the knowledge embedded within large language models (LLMs) without extensive retraining. However, current model editing…