3 papers
cs.LG2026
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
Zhiyuan Xu, Joseph Gardiner, Sana Belguith +1
Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capac…
cs.CR2025
Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
Zhiyuan Xu, Stanislav Abaimov, Joseph Gardiner +1
Modern large language models (LLMs) are typically secured by auditing data, prompts, and refusal policies, while treating the forward pass as an implementation detail. We show that…
cs.CR2025
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
Zhiyuan Xu, Joseph Gardiner, Sana Belguith
Large language models are typically trained on vast amounts of data during the pre-training phase, which may include some potentially harmful information. Fine-tuning attacks can e…