most citedThe dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2026

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner +2

Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, large…

cs.CL2026

Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters

Zhiyu Xu, Lean Wang, Yuanxin Liu +5

Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-sp…

cs.LG2026

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs

Zhiyuan Xu, Joseph Gardiner, Sana Belguith +1

Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capac…

stat.ME2026

Scalable Text-Embedding-informed Cognitive Diagnosis of Large Language Models

Jia Liu, Zhiyu Xu, Yuqi Gu

Large language models (LLMs) have achieved remarkable performance on diverse benchmarks, yet existing evaluation practices largely rely on coarse summary metrics that obscure under…

cs.CR2025

Progressive Behavioral Drift through Compression Valleys in Large Language Models

Zhiyuan Xu, Stanislav Abaimov, Joseph Gardiner +1

We show that attention sinks and compression valleys create a vulnerable region in decoder-only Transformers, where small activation perturbations can be amplified through the auto…

cs.CR2025★ 1 cited

The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models

Zhiyuan Xu, Joseph Gardiner, Sana Belguith

Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can e…