2 papers
cs.LG2025
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
Zirui He, Haiyan Zhao, Yiran Qiao +4
The ability of large language models (LLMs) to follow instructions is crucial for their practical applications, yet the underlying mechanisms remain poorly understood. This paper p…
cs.CL2024
Data-centric NLP Backdoor Defense from the Lens of Memorization
Zhenting Wang, Zhizhi Wang, Mingyu Jin +3
Backdoor attack is a severe threat to the trustworthiness of DNN-based language models. In this paper, we first extend the definition of memorization of language models from sample…