1 citations · 1 across the 8 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
Wei Zhao, Zhe Li, Peixin Zhang +1
Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this pro…
cs.AI2024
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Wei Zhao, Zhe Li, Yige Li +2
Large language models (LLMs) are increasingly being adopted in a wide range of real-world applications. Despite their impressive performance, recent studies have shown that LLMs ar…
cs.AI2023
Causality Analysis for Evaluating the Security of Large Language Models
Wei Zhao, Zhe Li, Jun Sun
Large Language Models (LLMs) such as GPT and Llama2 are increasingly adopted in many safety-critical applications. Their security is thus essential. Even with considerable efforts…