collaborators

5 papers

cs.CR2026

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Yizhe Zeng, Chenxu Niu, Wei Zhang +7

Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and cl…

cs.CL2026

LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language Models

Wei Zhang, Lintong Du, Yuanhe Zhang +4

Despite the strong performance of Large Language Models (LLMs) on complex instruction-following tasks, precise control of output length remains a persistent challenge. Existing met…

cs.CL2024

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings

Yuanhe Zhang, Zhenhong Zhou, Wei Zhang +4

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks yet still are vulnerable to external threats, particularly LLM Denial-of-Service (LLM-DoS…

cs.LG2024

Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models

Xiao Li, Wei Zhang, Zhuhong Li +6

Aligned Large Language Models (LLMs) have attracted significant attention for their safety, particularly in the context of jailbreak attacks that attempt to bypass guardrails via a…

cs.CL2024

Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions

Quan Liu, Zhenhong Zhou, Longzhu He +3

Large language models are susceptible to jailbreak attacks, which can result in the generation of harmful content. While prior defenses mitigate these risks by perturbing or inspec…