collaborators

5 papers

cs.CR2026

Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models

Jinman Wu, Yi Xie, Shen Lin +2

Safety alignment is often conceptualized as a monolithic process wherein harmfulness detection automatically triggers refusal. However, the persistence of jailbreak attacks suggest…

cs.CR2026

Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads

Jinman Wu, Yi Xie, Shiqian Zhao +1

Currently, open-sourced large language models (OSLLMs) have demonstrated remarkable generative performance. However, as their structure and weights are made public, they are expose…

cs.CR2026

WADBERT: Dual-channel Web Attack Detection Based on BERT Models

Kangqiang Luo, Yi Xie, Shiqian Zhao +1

Web attack detection is the first line of defense for securing web applications, designed to preemptively identify malicious activities. Deep learning-based approaches are increasi…

cs.CR2026

Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models

Shiqian Zhao, Chong Wang, Yiming Li +7

Text-to-Image (T2I) models, represented by DALLE and Midjourney, have gained huge popularity for creating realistic images. The quality of these images relies on the careful…

cs.CL2025

Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers

Viet-Anh Nguyen, Shiqian Zhao, Gia Dao +3

Recently, Large Reasoning Models (LRMs) have demonstrated superior logical capabilities compared to traditional Large Language Models (LLMs), gaining significant attention. Despite…