1 citations · 1 across the 3 of their papers we have counts for
3 papers
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner +2
Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, large…
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
Zhiyuan Xu, Joseph Gardiner, Sana Belguith +1
Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capac…
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
Zhiyuan Xu, Joseph Gardiner, Sana Belguith
Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can e…