activity
20232026
most citedSafety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions

2 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CR2026

KeyPooling: Measuring Where LLM API Relay Paths Collapse Prompt Cache Isolation

Bowen Sun, Yixi Cai, Xiaogeng Liu +3

Large language model (LLM) API relays authenticate customers separately but often forward requests through shared provider credentials. Providers scope prompt caches to upstream pr…

cs.CR2026

Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services

Bowen Sun, Zhengyue Zhao, Xiaogeng Liu +2

Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting…

cs.CR2026

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu +3

Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autor…

cs.CL2024★ 2 cited

Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions

Xiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi +3

With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignm…

cs.CV2023

Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion?

Zhengyue Zhao, Jinhao Duan, Kaidi Xu +5

Stable Diffusion has established itself as a foundation model in generative AI artistic applications, receiving widespread research and application. Some recent fine-tuning methods…