1 paper
Wei Zhao, Zhe Li, Peixin Zhang +1
Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this pro…