1 paper
Ying JinCheng, Minghui Xu, Yinhao Xiao +2
Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set…